Initial commit: local video captioning and translation tool
This commit is contained in:
+20
@@ -0,0 +1,20 @@
|
||||
# Virtual environment
|
||||
venv/
|
||||
.venv/
|
||||
|
||||
# Environment configurations and keys
|
||||
.env
|
||||
|
||||
# Python caching
|
||||
__pycache__/
|
||||
*.pyc
|
||||
|
||||
# Local media outputs and temporary files
|
||||
*.mp4
|
||||
*.wav
|
||||
*.srt
|
||||
temp_*
|
||||
*temp*
|
||||
|
||||
# macOS system files
|
||||
.DS_Store
|
||||
@@ -0,0 +1,118 @@
|
||||
# Video Captioning & Translation Tool (OpenAI API)
|
||||
|
||||
This is a lightweight Command Line Interface (CLI) tool that transcribes, translates, and captions video files. It uses **OpenAI's cloud APIs** for transcription (Whisper) and translation (GPT), combined with a **local FFmpeg pipeline** to extract audio and burn the translated subtitles into the output video.
|
||||
|
||||
## Architecture & Features
|
||||
|
||||
- **No Heavy Model Downloads**: Swapping local models for OpenAI APIs means zero local machine inference overhead and a tiny installation footprint (no PyTorch/CUDA/CoreML weights needed).
|
||||
- **Whisper API Transcription**: Calls OpenAI's `whisper-1` model with `verbose_json` to fetch accurate segment timestamps.
|
||||
- **Context-Aware Translation**: Translates transcription segments with a rolling context window (remembers prior segments) via the OpenAI Chat Completions API (`gpt-4o-mini` by default). This maintains high translation cohesion.
|
||||
- **Local Subtitle Rendering**: Uses local FFmpeg filters to hardcode (burn) subtitles into the output video.
|
||||
|
||||
---
|
||||
|
||||
## Dependencies & Setup
|
||||
|
||||
### 1. System Requirements
|
||||
You must have **FFmpeg** installed locally to handle audio extraction and video rendering.
|
||||
On macOS, install it using Homebrew:
|
||||
```bash
|
||||
brew install ffmpeg
|
||||
```
|
||||
|
||||
### 2. Python Environment Setup
|
||||
Create a virtual environment and install the dependencies:
|
||||
```bash
|
||||
# Navigate to the project folder
|
||||
cd /Users/adolforeyna/.gemini/antigravity/scratch/local-translator
|
||||
|
||||
# Create a virtual environment
|
||||
python3 -m venv venv
|
||||
source venv/bin/activate
|
||||
|
||||
# Install requirements
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### 3. API Configuration
|
||||
The tool reads `.env` configuration. A `.env` file has been automatically copied from your `pythonwhisper` project containing your `OPENAI_API_KEY`.
|
||||
|
||||
If you need to update it, modify the `.env` file in the project folder:
|
||||
```env
|
||||
OPENAI_API_KEY=your_openai_api_key_here
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Usage Examples
|
||||
|
||||
Always activate your virtual environment before running the tool:
|
||||
```bash
|
||||
source venv/bin/activate
|
||||
```
|
||||
|
||||
### 1. Translate English Video to Spanish (Default)
|
||||
```bash
|
||||
python main.py -v /path/to/my_video.mp4
|
||||
```
|
||||
*Outputs: `/path/to/my_video_es_captioned.mp4`*
|
||||
|
||||
### 2. Translate Spanish Video to English
|
||||
```bash
|
||||
python main.py -v /path/to/video.mp4 -s es -t en
|
||||
```
|
||||
|
||||
### 3. Keep Temporary Subtitles and Audio Files
|
||||
If you want to keep the intermediate `.srt` subtitle file and `.wav` extracted audio for manual editing or verification, add `--keep-temp`:
|
||||
```bash
|
||||
python main.py -v /path/to/video.mp4 --keep-temp
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Visuals & Style Customization
|
||||
|
||||
The tool supports subtitle visual customization via **FFmpeg ASS Style Presets** (available when burning subtitles):
|
||||
- **`box`** (Default): White text on a semi-transparent dark background box. Offers the best readability.
|
||||
- **`default`**: Standard white text with outline and drop shadow.
|
||||
- **`yellow`**: Yellow text with black outline and drop shadow.
|
||||
- **`clean`**: Larger white text on a minimal thin outline with a Trebuchet font style.
|
||||
|
||||
To select a preset:
|
||||
```bash
|
||||
python main.py -v my_video.mp4 --style yellow
|
||||
```
|
||||
|
||||
For advanced styling, pass a raw ASS style string using `--custom-style`:
|
||||
```bash
|
||||
python main.py -v my_video.mp4 --custom-style "FontName=Arial,FontSize=24,PrimaryColour=&H0000FFFF,BorderStyle=3,BackColour=&H80000000"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Cadence Customization (Segment Splitting)
|
||||
|
||||
By default, the translation tool preserves the transcription segment lengths returned by the Whisper API. If you have long sentences or fast spoken parts, subtitles can get cluttered.
|
||||
|
||||
You can enforce a **character limit per card** using the `--max-chars` flag. The tool will automatically split long sentences into smaller, logically divided segments and distribute the timing proportionally:
|
||||
```bash
|
||||
# Split subtitles so each card is at most 40 characters
|
||||
python main.py -v my_video.mp4 --max-chars 40
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## CLI Argument Reference
|
||||
|
||||
| Argument | Short | Description |
|
||||
| :--- | :--- | :--- |
|
||||
| `--video` | `-v` | **Required**. Path to the input video file. |
|
||||
| `--source-lang` | `-s` | Source language of the video (default: `en`). |
|
||||
| `--target-lang` | `-t` | Target language for the captions (default: `es`). |
|
||||
| `--output` | `-o` | Custom path for the output captioned video. |
|
||||
| `--whisper-model`| | OpenAI Whisper model to use (default: `whisper-1`). |
|
||||
| `--model` | | OpenAI translation model to use (default: `gpt-4o-mini`). |
|
||||
| `--keep-temp` | | Keep intermediate `.wav` and `.srt` files. |
|
||||
| `--style` | | Subtitle style preset: `default`, `yellow`, `box`, `clean` (default: `box`). |
|
||||
| `--custom-style` | | Raw ASS style string override (e.g. `FontName=Arial,FontSize=24`). |
|
||||
| `--max-chars` | | Force subtitle line length limit by splitting segments (default: `0` for no splitting). |
|
||||
@@ -0,0 +1,41 @@
|
||||
import os
|
||||
from dotenv import load_dotenv
|
||||
|
||||
# Load environment variables from .env
|
||||
load_dotenv()
|
||||
|
||||
# API Keys and tokens
|
||||
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
|
||||
|
||||
# Whisper ASR Settings
|
||||
DEFAULT_WHISPER_MODEL = "whisper-1"
|
||||
|
||||
# LLM Translation settings
|
||||
DEFAULT_OPENAI_MODEL = "gpt-4o-mini"
|
||||
|
||||
# Defaults
|
||||
DEFAULT_SOURCE_LANG = "en"
|
||||
DEFAULT_TARGET_LANG = "es"
|
||||
|
||||
# Target translation languages map
|
||||
SUPPORTED_LANGUAGES = {
|
||||
"en": "English",
|
||||
"es": "Spanish",
|
||||
"fr": "French",
|
||||
"de": "German",
|
||||
"it": "Italian",
|
||||
"pt": "Portuguese",
|
||||
"zh": "Chinese",
|
||||
"ja": "Japanese",
|
||||
"ar": "Arabic",
|
||||
"ru": "Russian"
|
||||
}
|
||||
|
||||
# Subtitle Visual Presets (ASS style rules for FFmpeg)
|
||||
SUBTITLE_PRESETS = {
|
||||
"default": "FontName=Helvetica,FontSize=20,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,BorderStyle=1,Outline=1,Shadow=1,MarginV=15",
|
||||
"yellow": "FontName=Helvetica,FontSize=22,PrimaryColour=&H0000FFFF,OutlineColour=&H00000000,BorderStyle=1,Outline=1.5,Shadow=1,MarginV=15",
|
||||
"box": "FontName=Helvetica,FontSize=20,PrimaryColour=&H00FFFFFF,BackColour=&H80000000,BorderStyle=3,Outline=0,Shadow=0,MarginV=15",
|
||||
"clean": "FontName=Arial,FontSize=22,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,BorderStyle=1,Outline=0.5,Shadow=0,MarginV=20"
|
||||
}
|
||||
DEFAULT_STYLE = "box"
|
||||
@@ -0,0 +1,171 @@
|
||||
import os
|
||||
import sys
|
||||
import argparse
|
||||
import time
|
||||
import shutil
|
||||
import config
|
||||
from translator import extract_audio, transcribe_audio, translate_segments, write_srt, burn_subtitles
|
||||
|
||||
def check_dependencies():
|
||||
"""Verify that FFmpeg is installed and accessible in system path"""
|
||||
if not shutil.which("ffmpeg"):
|
||||
print("[Error] FFmpeg is not installed or not in your system PATH.")
|
||||
print("Please install FFmpeg to run this tool. On macOS: 'brew install ffmpeg'")
|
||||
sys.exit(1)
|
||||
|
||||
def main():
|
||||
check_dependencies()
|
||||
|
||||
parser = argparse.ArgumentParser(description="Video Captioning & Translation Tool using OpenAI API + Local FFmpeg")
|
||||
|
||||
parser.add_argument("-v", "--video", type=str, required=True, help="Path to the input video file")
|
||||
parser.add_argument("-s", "--source-lang", type=str, default=config.DEFAULT_SOURCE_LANG,
|
||||
help=f"Source language of the video audio (default: {config.DEFAULT_SOURCE_LANG})")
|
||||
parser.add_argument("-t", "--target-lang", type=str, default=config.DEFAULT_TARGET_LANG,
|
||||
help=f"Target language for the captions (default: {config.DEFAULT_TARGET_LANG})")
|
||||
parser.add_argument("-o", "--output", type=str, default=None,
|
||||
help="Path for the output captioned video (default: <video_name>_<target>_captioned.mp4)")
|
||||
parser.add_argument("--whisper-model", type=str, default=config.DEFAULT_WHISPER_MODEL,
|
||||
help=f"OpenAI Whisper model name (default: {config.DEFAULT_WHISPER_MODEL})")
|
||||
parser.add_argument("--model", type=str, default=config.DEFAULT_OPENAI_MODEL,
|
||||
help=f"OpenAI GPT model name for translation (default: {config.DEFAULT_OPENAI_MODEL})")
|
||||
parser.add_argument("--keep-temp", action="store_true", default=False,
|
||||
help="Keep intermediate audio (.wav) and subtitle (.srt) files")
|
||||
parser.add_argument("--style", type=str, choices=list(config.SUBTITLE_PRESETS.keys()), default=config.DEFAULT_STYLE,
|
||||
help=f"Subtitle style preset: {', '.join(config.SUBTITLE_PRESETS.keys())} (default: {config.DEFAULT_STYLE})")
|
||||
parser.add_argument("--custom-style", type=str, default=None,
|
||||
help="Raw ASS style override (e.g. 'FontName=Arial,FontSize=24,PrimaryColour=&H0000FFFF')")
|
||||
parser.add_argument("--max-chars", type=int, default=0,
|
||||
help="Force subtitle line length limit by splitting segments (default: 0 for no splitting)")
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
video_path = os.path.abspath(args.video)
|
||||
if not os.path.exists(video_path):
|
||||
print(f"[Error] Input video file does not exist: {args.video}")
|
||||
sys.exit(1)
|
||||
|
||||
# Define output file path if not specified
|
||||
video_dir = os.path.dirname(video_path)
|
||||
video_name, video_ext = os.path.splitext(os.path.basename(video_path))
|
||||
|
||||
if args.output:
|
||||
output_path = os.path.abspath(args.output)
|
||||
else:
|
||||
output_path = os.path.join(video_dir, f"{video_name}_{args.target_lang}_captioned{video_ext}")
|
||||
|
||||
# Use space-free temporary names to prevent FFmpeg subtitle filter parsing errors
|
||||
temp_audio_path = os.path.join(video_dir, f"temp_audio_{int(time.time())}.wav")
|
||||
temp_srt_path = os.path.join(video_dir, f"temp_subtitles_{int(time.time())}.srt")
|
||||
|
||||
# User-friendly SRT path for when keep-temp is enabled
|
||||
final_srt_path = os.path.join(video_dir, f"{video_name}_{args.target_lang}.srt")
|
||||
|
||||
print("\n" + "="*50)
|
||||
print(" VIDEO CAPTIONING & TRANSLATION TOOL (OPENAI)")
|
||||
print("="*50)
|
||||
print(f"Input Video: {video_path}")
|
||||
print(f"Source Language: {args.source_lang}")
|
||||
print(f"Target Language: {args.target_lang}")
|
||||
print(f"Output Video: {output_path}")
|
||||
print(f"Whisper Model: {args.whisper_model}")
|
||||
print(f"Translation Model:{args.model}")
|
||||
print("="*50 + "\n")
|
||||
|
||||
pipeline_start = time.time()
|
||||
|
||||
# Step 1: Extract Audio
|
||||
step_start = time.time()
|
||||
if not extract_audio(video_path, temp_audio_path):
|
||||
print("[Error] Failed to extract audio. Exiting.")
|
||||
sys.exit(1)
|
||||
audio_time = time.time() - step_start
|
||||
|
||||
# Step 2: Transcribe Audio via OpenAI API
|
||||
step_start = time.time()
|
||||
try:
|
||||
segments = transcribe_audio(temp_audio_path, args.whisper_model)
|
||||
except Exception as e:
|
||||
print(f"[Error] ASR failed: {e}")
|
||||
if not args.keep_temp and os.path.exists(temp_audio_path):
|
||||
os.remove(temp_audio_path)
|
||||
sys.exit(1)
|
||||
transcribe_time = time.time() - step_start
|
||||
|
||||
if not segments:
|
||||
print("[Warning] No speech detected in video audio.")
|
||||
shutil.copy2(video_path, output_path)
|
||||
print(f"[Finished] Output copied to {output_path}")
|
||||
if not args.keep_temp and os.path.exists(temp_audio_path):
|
||||
os.remove(temp_audio_path)
|
||||
sys.exit(0)
|
||||
|
||||
# Step 3: Translate segments via OpenAI API
|
||||
step_start = time.time()
|
||||
try:
|
||||
translated_segments = translate_segments(
|
||||
segments=segments,
|
||||
source_lang=args.source_lang,
|
||||
target_lang=args.target_lang,
|
||||
model=args.model
|
||||
)
|
||||
except Exception as e:
|
||||
print(f"[Error] Translation failed: {e}")
|
||||
if not args.keep_temp and os.path.exists(temp_audio_path):
|
||||
os.remove(temp_audio_path)
|
||||
sys.exit(1)
|
||||
translation_time = time.time() - step_start
|
||||
|
||||
# Optional: Adjust subtitle cadence by splitting long segments
|
||||
if args.max_chars > 0:
|
||||
from translator import split_long_segments
|
||||
translated_segments = split_long_segments(translated_segments, args.max_chars)
|
||||
|
||||
# Step 4: Write Subtitles
|
||||
write_srt(translated_segments, temp_srt_path)
|
||||
|
||||
# Step 5: Burn Subtitles locally using FFmpeg
|
||||
step_start = time.time()
|
||||
burn_success = burn_subtitles(
|
||||
video_path=video_path,
|
||||
srt_path=temp_srt_path,
|
||||
output_path=output_path,
|
||||
style_preset=args.style,
|
||||
custom_style=args.custom_style
|
||||
)
|
||||
burn_time = time.time() - step_start
|
||||
|
||||
# Clean up temp files
|
||||
if args.keep_temp:
|
||||
friendly_audio_path = os.path.join(video_dir, f"{video_name}_temp_audio.wav")
|
||||
if os.path.exists(temp_audio_path):
|
||||
shutil.move(temp_audio_path, friendly_audio_path)
|
||||
print(f"[Cleanup] Saved audio track copy at: {friendly_audio_path}")
|
||||
if os.path.exists(temp_srt_path):
|
||||
shutil.move(temp_srt_path, final_srt_path)
|
||||
print(f"[Cleanup] Saved subtitle file copy at: {final_srt_path}")
|
||||
else:
|
||||
print("[Cleanup] Removing temporary files...")
|
||||
if os.path.exists(temp_audio_path):
|
||||
os.remove(temp_audio_path)
|
||||
if os.path.exists(temp_srt_path):
|
||||
os.remove(temp_srt_path)
|
||||
|
||||
if burn_success:
|
||||
print("\n" + "="*50)
|
||||
print(f"SUCCESS: Captioned video generated!")
|
||||
print(f"Saved at: {output_path}")
|
||||
print("-"*50)
|
||||
print(f"Timing Breakdown:")
|
||||
print(f" - Audio Extraction: {audio_time:.1f}s")
|
||||
print(f" - OpenAI ASR: {transcribe_time:.1f}s")
|
||||
print(f" - OpenAI Translate: {translation_time:.1f}s")
|
||||
print(f" - FFmpeg Burning: {burn_time:.1f}s")
|
||||
print(f" - Total Elapsed: {time.time() - pipeline_start:.1f}s")
|
||||
print("="*50 + "\n")
|
||||
else:
|
||||
print("\n[Error] Failed to generate captioned video during subtitle burning phase.")
|
||||
sys.exit(1)
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,4 @@
|
||||
openai
|
||||
python-dotenv
|
||||
tqdm
|
||||
requests
|
||||
@@ -0,0 +1,68 @@
|
||||
import os
|
||||
import sys
|
||||
import shutil
|
||||
|
||||
def run_diagnostics():
|
||||
print("="*50)
|
||||
print(" DIAGNOSTICS & SYSTEM CHECKS (OPENAI)")
|
||||
print("="*50)
|
||||
|
||||
# 1. Check Python Dependencies
|
||||
print("[1] Checking python imports...")
|
||||
deps = ["openai", "dotenv", "tqdm", "requests"]
|
||||
for dep in deps:
|
||||
try:
|
||||
__import__(dep)
|
||||
print(f" - {dep}: OK")
|
||||
except ImportError:
|
||||
print(f" - {dep}: NOT INSTALLED (run: pip install {dep})")
|
||||
|
||||
# 2. Check FFmpeg
|
||||
print("\n[2] Checking FFmpeg...")
|
||||
ffmpeg_path = shutil.which("ffmpeg")
|
||||
if ffmpeg_path:
|
||||
print(f" - FFmpeg found: {ffmpeg_path}")
|
||||
else:
|
||||
print(" - FFmpeg NOT found in system PATH. Subtitle burning will fail.")
|
||||
print(" Install it via: brew install ffmpeg")
|
||||
|
||||
# 3. Check Configurations
|
||||
print("\n[3] Checking configurations & environment...")
|
||||
try:
|
||||
import config
|
||||
print(f" - Config loaded successfully.")
|
||||
print(f" - OPENAI_API_KEY: {'Configured' if config.OPENAI_API_KEY else 'Not Configured (Must set in .env)'}")
|
||||
print(f" - Default Whisper model: {config.DEFAULT_WHISPER_MODEL}")
|
||||
print(f" - Default GPT model: {config.DEFAULT_OPENAI_MODEL}")
|
||||
except Exception as e:
|
||||
print(f" - Failed to load config: {e}")
|
||||
|
||||
# 4. Test OpenAI Connection
|
||||
print("\n[4] Testing OpenAI Connection...")
|
||||
if not config.OPENAI_API_KEY:
|
||||
print(" - OpenAI Test: SKIPPED (API Key missing)")
|
||||
print("="*50)
|
||||
return
|
||||
|
||||
try:
|
||||
from translator import OpenAIAPITranslator
|
||||
|
||||
translator = OpenAIAPITranslator()
|
||||
test_text = "Hello, how are you today?"
|
||||
|
||||
print(f" - Translating phrase: \"{test_text}\" to Spanish...")
|
||||
translated = translator.translate_text(test_text, "en", "es")
|
||||
|
||||
print(f" - Output: \"{translated}\"")
|
||||
if translated != test_text:
|
||||
print(" - OpenAI Connection Test: SUCCESS")
|
||||
else:
|
||||
print(" - OpenAI Connection Test: WARNING (returned original text, check API usage)")
|
||||
except Exception as e:
|
||||
print(f" - OpenAI Connection Test: FAILED ({e})")
|
||||
print(" Please verify your OPENAI_API_KEY in the .env file.")
|
||||
|
||||
print("\n" + "="*50)
|
||||
|
||||
if __name__ == "__main__":
|
||||
run_diagnostics()
|
||||
+356
@@ -0,0 +1,356 @@
|
||||
import os
|
||||
import subprocess
|
||||
import time
|
||||
from tqdm import tqdm
|
||||
from openai import OpenAI
|
||||
import config
|
||||
|
||||
def format_timestamp(seconds):
|
||||
"""Convert float seconds to SRT timestamp format: HH:MM:SS,mmm"""
|
||||
hours = int(seconds // 3600)
|
||||
minutes = int((seconds % 3600) // 60)
|
||||
secs = int(seconds % 60)
|
||||
milliseconds = int(round((seconds % 1) * 1000))
|
||||
if milliseconds == 1000:
|
||||
secs += 1
|
||||
milliseconds = 0
|
||||
return f"{hours:02d}:{minutes:02d}:{secs:02d},{milliseconds:03d}"
|
||||
|
||||
def extract_audio(video_path, audio_path):
|
||||
"""Extract mono 16kHz audio from a video using FFmpeg"""
|
||||
print(f"[Audio] Extracting audio from {os.path.basename(video_path)}...")
|
||||
cmd = [
|
||||
"ffmpeg", "-y",
|
||||
"-i", video_path,
|
||||
"-ar", "16000",
|
||||
"-ac", "1",
|
||||
"-c:a", "pcm_s16le",
|
||||
audio_path
|
||||
]
|
||||
try:
|
||||
subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True, check=True)
|
||||
print("[Audio] Audio extraction complete.")
|
||||
return True
|
||||
except subprocess.CalledProcessError as e:
|
||||
print(f"[Audio Error] FFmpeg extraction failed: {e.stderr}")
|
||||
return False
|
||||
|
||||
def transcribe_audio(audio_path, model_name=config.DEFAULT_WHISPER_MODEL):
|
||||
"""Transcribe audio file using OpenAI's Whisper API and return segments with timestamps"""
|
||||
if not config.OPENAI_API_KEY:
|
||||
raise ValueError("OPENAI_API_KEY is not configured in .env file.")
|
||||
|
||||
print(f"[ASR] Sending audio to OpenAI Whisper API using model '{model_name}'...")
|
||||
client = OpenAI(api_key=config.OPENAI_API_KEY)
|
||||
|
||||
start_time = time.time()
|
||||
with open(audio_path, "rb") as audio_file:
|
||||
transcript = client.audio.transcriptions.create(
|
||||
model=model_name,
|
||||
file=audio_file,
|
||||
response_format="verbose_json"
|
||||
)
|
||||
duration = time.time() - start_time
|
||||
print(f"[ASR] Transcription complete in {duration:.2f} seconds.")
|
||||
|
||||
# Extract segments
|
||||
raw_segments = []
|
||||
if hasattr(transcript, "segments"):
|
||||
raw_segments = transcript.segments
|
||||
elif isinstance(transcript, dict):
|
||||
raw_segments = transcript.get("segments", [])
|
||||
else:
|
||||
raw_segments = getattr(transcript, "segments", [])
|
||||
|
||||
segments = []
|
||||
for seg in raw_segments:
|
||||
if isinstance(seg, dict):
|
||||
segments.append({
|
||||
"start": seg.get("start", 0.0),
|
||||
"end": seg.get("end", 0.0),
|
||||
"text": seg.get("text", "").strip()
|
||||
})
|
||||
else:
|
||||
segments.append({
|
||||
"start": getattr(seg, "start", 0.0),
|
||||
"end": getattr(seg, "end", 0.0),
|
||||
"text": getattr(seg, "text", "").strip()
|
||||
})
|
||||
|
||||
print(f"[ASR] Transcribed {len(segments)} segments.")
|
||||
return segments
|
||||
|
||||
class OpenAIAPITranslator:
|
||||
def __init__(self, model=config.DEFAULT_OPENAI_MODEL, api_key=None):
|
||||
key = api_key or config.OPENAI_API_KEY
|
||||
if not key:
|
||||
raise ValueError("OpenAI API key must be provided or set in environment.")
|
||||
self.client = OpenAI(api_key=key)
|
||||
self.model = model
|
||||
|
||||
def translate_text(self, text, source_lang, target_lang, context_history=None):
|
||||
"""Translate a single block of text using OpenAI GPT model, with optional context history"""
|
||||
source_name = config.SUPPORTED_LANGUAGES.get(source_lang, source_lang)
|
||||
target_name = config.SUPPORTED_LANGUAGES.get(target_lang, target_lang)
|
||||
|
||||
system_prompt = (
|
||||
f"You are an expert, professional subtitle translator.\n"
|
||||
f"Your task is to translate the given text from {source_name} to {target_name}.\n\n"
|
||||
f"Hard Rules:\n"
|
||||
f"1. Preserve the meaning, tone, and slang style accurately.\n"
|
||||
f"2. Output ONLY the translated text. Do not include any notes, explanations, introductory/outro remarks, quotes, or formatting.\n"
|
||||
f"3. Make it natural and fitting for screen subtitles (short, punchy sentences where appropriate).\n"
|
||||
f"4. If you cannot translate it (e.g. it's nonsensical), output the original text unchanged."
|
||||
)
|
||||
|
||||
user_content = ""
|
||||
if context_history:
|
||||
user_content += "--- CONTEXT ---\n"
|
||||
for prev_src, prev_trans in context_history:
|
||||
user_content += f"Previous original: \"{prev_src}\"\n"
|
||||
user_content += f"Previous translation: \"{prev_trans}\"\n\n"
|
||||
user_content += "--- CURRENT SEGMENT TO TRANSLATE ---\n"
|
||||
|
||||
user_content += f"Original text: \"{text}\"\nTranslation:"
|
||||
|
||||
messages = [
|
||||
{"role": "system", "content": system_prompt},
|
||||
{"role": "user", "content": user_content}
|
||||
]
|
||||
|
||||
try:
|
||||
response = self.client.chat.completions.create(
|
||||
model=self.model,
|
||||
messages=messages,
|
||||
temperature=0.3,
|
||||
max_tokens=256
|
||||
)
|
||||
translated = response.choices[0].message.content.strip()
|
||||
# Remove surrounding quotes if model added them
|
||||
if translated.startswith('"') and translated.endswith('"'):
|
||||
translated = translated[1:-1]
|
||||
return translated
|
||||
except Exception as e:
|
||||
print(f"\n[OpenAI Error] Translation failed: {e}")
|
||||
return text
|
||||
|
||||
def translate_segments(segments, source_lang, target_lang, model=config.DEFAULT_OPENAI_MODEL):
|
||||
"""Translate Whisper segments list using OpenAIAPITranslator with a rolling context window"""
|
||||
print(f"[Translation] Translating segments using OpenAI model '{model}' from '{source_lang}' to '{target_lang}'...")
|
||||
translator = OpenAIAPITranslator(model=model)
|
||||
|
||||
translated_segments = []
|
||||
context_window = [] # Keep a rolling window of (original, translation) pairs (max 2)
|
||||
|
||||
for segment in tqdm(segments, desc="Translating", unit="segment"):
|
||||
text = segment.get("text", "").strip()
|
||||
if not text:
|
||||
translated_segments.append({
|
||||
"start": segment["start"],
|
||||
"end": segment["end"],
|
||||
"text": ""
|
||||
})
|
||||
continue
|
||||
|
||||
translated_text = translator.translate_text(
|
||||
text=text,
|
||||
source_lang=source_lang,
|
||||
target_lang=target_lang,
|
||||
context_history=context_window
|
||||
)
|
||||
|
||||
translated_segments.append({
|
||||
"start": segment["start"],
|
||||
"end": segment["end"],
|
||||
"text": translated_text
|
||||
})
|
||||
|
||||
# Update rolling context
|
||||
context_window.append((text, translated_text))
|
||||
if len(context_window) > 2:
|
||||
context_window.pop(0)
|
||||
|
||||
print(f"[Translation] Completed translation of {len(translated_segments)} segments.")
|
||||
return translated_segments
|
||||
|
||||
def split_long_segments(segments, max_chars=40):
|
||||
"""Split segments that exceed max_chars into smaller, proportionally-timed segments"""
|
||||
if not max_chars or max_chars <= 0:
|
||||
return segments
|
||||
|
||||
refined = []
|
||||
for seg in segments:
|
||||
text = seg["text"].strip()
|
||||
start = seg["start"]
|
||||
end = seg["end"]
|
||||
duration = end - start
|
||||
|
||||
if len(text) <= max_chars or duration <= 0:
|
||||
refined.append(seg)
|
||||
continue
|
||||
|
||||
words = text.split()
|
||||
chunks = []
|
||||
current_chunk = []
|
||||
current_len = 0
|
||||
|
||||
for word in words:
|
||||
word_len = len(word)
|
||||
added_len = word_len + (1 if current_len > 0 else 0)
|
||||
|
||||
if current_len + added_len > max_chars and current_chunk:
|
||||
chunks.append(" ".join(current_chunk))
|
||||
current_chunk = [word]
|
||||
current_len = word_len
|
||||
else:
|
||||
current_chunk.append(word)
|
||||
current_len += added_len
|
||||
|
||||
if current_chunk:
|
||||
chunks.append(" ".join(current_chunk))
|
||||
|
||||
# Distribute timing proportionally
|
||||
total_chars = sum(len(c) for c in chunks)
|
||||
if total_chars == 0:
|
||||
refined.append(seg)
|
||||
continue
|
||||
|
||||
current_time = start
|
||||
for chunk in chunks:
|
||||
chunk_len = len(chunk)
|
||||
chunk_dur = duration * (chunk_len / total_chars)
|
||||
refined.append({
|
||||
"start": current_time,
|
||||
"end": current_time + chunk_dur,
|
||||
"text": chunk
|
||||
})
|
||||
current_time += chunk_dur
|
||||
|
||||
print(f"[Cadence] Split {len(segments)} segments into {len(refined)} segments (max_chars={max_chars}).")
|
||||
return refined
|
||||
|
||||
def write_srt(segments, srt_path):
|
||||
"""Write subtitle segments to an SRT file"""
|
||||
print(f"[Subtitles] Writing subtitles to {srt_path}...")
|
||||
with open(srt_path, "w", encoding="utf-8") as f:
|
||||
for idx, segment in enumerate(segments, 1):
|
||||
start_str = format_timestamp(segment["start"])
|
||||
end_str = format_timestamp(segment["end"])
|
||||
text = segment["text"].replace("\n", " ").strip()
|
||||
|
||||
f.write(f"{idx}\n")
|
||||
f.write(f"{start_str} --> {end_str}\n")
|
||||
f.write(f"{text}\n\n")
|
||||
print("[Subtitles] SRT file written.")
|
||||
|
||||
def check_subtitles_filter_support():
|
||||
"""Check if the subtitles filter is available in FFmpeg"""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["ffmpeg", "-filters"],
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
text=True
|
||||
)
|
||||
return " subtitles " in result.stdout
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
def embed_subtitles_soft(video_path, srt_path, output_path):
|
||||
"""Embed subtitles as a soft track in the video using mov_text codec (requires no filter/libass)"""
|
||||
print(f"[Video] Embedding subtitles as soft track into video...")
|
||||
|
||||
video_dir = os.path.dirname(os.path.abspath(video_path))
|
||||
srt_name = os.path.basename(srt_path)
|
||||
video_name = os.path.basename(video_path)
|
||||
output_name = os.path.basename(output_path)
|
||||
|
||||
cmd = [
|
||||
"ffmpeg", "-y",
|
||||
"-i", video_name,
|
||||
"-i", srt_name,
|
||||
"-c:v", "copy",
|
||||
"-c:a", "copy",
|
||||
"-c:s", "mov_text",
|
||||
output_name
|
||||
]
|
||||
|
||||
try:
|
||||
subprocess.run(
|
||||
cmd,
|
||||
cwd=video_dir,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
text=True,
|
||||
check=True
|
||||
)
|
||||
print(f"[Video] Subtitles successfully soft-embedded into {output_path}")
|
||||
return True
|
||||
except subprocess.CalledProcessError as e:
|
||||
print(f"[Video Error] Soft-embedding subtitles failed: {e.stderr}")
|
||||
return False
|
||||
|
||||
def burn_subtitles(video_path, srt_path, output_path, style_preset="box", custom_style=None):
|
||||
"""Adds subtitles to the video, either burning them (if supported) or softcoding them (as fallback)"""
|
||||
if not check_subtitles_filter_support():
|
||||
print("\n" + "!"*60)
|
||||
print("[Warning] Your current FFmpeg build does not support hardcoding (burning) subtitles.")
|
||||
print(" It is missing the 'subtitles' video filter (usually due to a missing libass library).")
|
||||
print(" We are falling back to soft-coding (embedding a subtitle track) instead.")
|
||||
print(" To fix this and enable hardcoded/burned captions, run:")
|
||||
print(" brew reinstall ffmpeg")
|
||||
print("!"*60 + "\n")
|
||||
return embed_subtitles_soft(video_path, srt_path, output_path)
|
||||
|
||||
print(f"[Video] Burning subtitles from {os.path.basename(srt_path)} into video...")
|
||||
|
||||
video_dir = os.path.dirname(os.path.abspath(video_path))
|
||||
srt_name = os.path.basename(srt_path)
|
||||
video_name = os.path.basename(video_path)
|
||||
output_name = os.path.basename(output_path)
|
||||
|
||||
# Resolve style string
|
||||
style_str = config.SUBTITLE_PRESETS.get(style_preset, config.SUBTITLE_PRESETS["box"])
|
||||
if custom_style:
|
||||
style_str = custom_style
|
||||
|
||||
print(f"[Video] Using style settings: {style_str}")
|
||||
|
||||
cmd = [
|
||||
"ffmpeg", "-y",
|
||||
"-i", video_name,
|
||||
"-vf", f"subtitles={srt_name}:force_style='{style_str}'",
|
||||
"-c:a", "copy", # Copy audio stream without re-encoding
|
||||
output_name
|
||||
]
|
||||
|
||||
try:
|
||||
print(f"[Video] Launching FFmpeg subtitle filter in directory: {video_dir}")
|
||||
subprocess.run(
|
||||
cmd,
|
||||
cwd=video_dir,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
text=True,
|
||||
check=True
|
||||
)
|
||||
print(f"[Video] Subtitles successfully burned into {output_path}")
|
||||
return True
|
||||
except subprocess.CalledProcessError as e:
|
||||
print(f"[Video Error] FFmpeg subtitle burning failed: {e.stderr}")
|
||||
print("[Video] Attempting absolute path fallback...")
|
||||
escaped_srt = os.path.abspath(srt_path).replace(":", "\\:").replace("'", "'\\''")
|
||||
fallback_cmd = [
|
||||
"ffmpeg", "-y",
|
||||
"-i", video_path,
|
||||
"-vf", f"subtitles='{escaped_srt}':force_style='{style_str}'",
|
||||
"-c:a", "copy",
|
||||
output_path
|
||||
]
|
||||
try:
|
||||
subprocess.run(fallback_cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True, check=True)
|
||||
print(f"[Video] Subtitles successfully burned into {output_path} (via absolute path)")
|
||||
return True
|
||||
except subprocess.CalledProcessError as err:
|
||||
print(f"[Video Error] Fallback FFmpeg failed: {err.stderr}")
|
||||
return False
|
||||
Reference in New Issue
Block a user