© PROJECT DETAIL
Offline-first multilingual video dubbing pipeline
AI-powered, offline-first multilingual video dubbing pipeline that automatically downloads, transcribes, translates, and synchronizes YouTube videos into English while preserving the original speaker’s voice.
© 01 / OVERVIEW
DubFlow is an AI-powered, offline-first multilingual video dubbing pipeline that automatically downloads, transcribes, translates, and synchronizes YouTube videos into English. It uses zero-shot voice cloning to preserve the original speaker's vocal characteristics while keeping the workflow modular and local.
DubFlow runs as a modular, job-based pipeline with isolated stages for downloading, extracting audio, transcribing, translating, synthesizing English speech, and muxing the final video.
© 02 / TECH STACK
© 03 / ARCHITECTURE
The pipeline uses yt-dlp for downloads, FFmpeg for audio extraction and final muxing, Faster-Whisper for transcription, NLLB-200 for translation, and Coqui XTTS-v2 for zero-shot voice cloning with short reference clips.
© 04 / CHALLENGES
Keeping each pipeline stage isolated so jobs can be retried and debugged independently
Preserving speaker identity while generating English speech from short reference audio clips
Keeping transcription, translation, and synthesis synchronized so the final video stays in time
Handling multilingual content with a modular workflow that stays offline-first
© 05 / IMPACT
DubFlow removes the need for cloud dubbing services by keeping the workflow local, modular, and privacy-preserving.