I Ran OmniVoice Again, and Hit Some Weird Errors
This time I dubbed an English video into Korean. I timed each step of a single dub, and sometimes the voice turned into weird robot noise.
So after the last post, I messed around with OmniVoice some more. This time I was curious about two things. One was how long it actually takes to dub a single video. The other was the opposite direction from last time: what happens if I turn an English video into Korean.
So I grabbed a short clip of a Trump speech (in English) and dubbed it into Korean. Here’s the original first:
Original: the English clip I wanted to dub into Korean
And here’s the result after OmniVoice dubbed it into Korean. It cloned the original voice and had it speak Korean:
OmniVoice: dubbed English → Korean, with the original voice cloned
How long does one dub take
Taking a single 22-second clip all the way from transcription to translation to voice synthesis to export took about 3 minutes total. All of it ran on my MacBook, no internet. Here’s how it breaks down step by step:
- Prep (pulling the audio out of the video and splitting voice from background): about 7 seconds
- Transcription (turning the speech into text): about 29 seconds
- Translation (English to Korean): about 90 seconds
- Building the voice profile (registering the original voice): about 5 seconds
- Voice synthesis + cloning: about 49 seconds
- Export (merging it back into the video): about 2 seconds
One fun detail: the first synthesis run takes longer, but running it again cuts the time roughly in half. That’s because the time it takes to load the AI model into memory only gets counted on that first run.
Which model ran at each step
The dub is split into steps, and a different model handles each one:
- Splitting voice and background: Demucs
- Transcription: WhisperX
- Word timing: wav2vec2
- Speaker separation (telling apart who’s talking): WavLM
- Translation: gemma2:27b (better quality than the built-in translator)
- Voice synthesis + cloning: OmniVoice
It wasn’t all smooth
A couple of things tripped me up along the way.
One, sometimes where a voice should’ve been, I got this crushed, staticky noise instead of an actual human voice. This time I had it build Korean from an English voice sample, and that voice trying to imitate Korean, a language it had never spoken, sometimes came out broken. So I switched to a setting that runs the voice a few more times to clean it up, re-ran it, and the Trump video came out fine.
Two, when I ran the translation, one sentence came out totally different from the original, so I had to go in and fix it by hand.
So, the takeaway
Of all the open-source dubbing tools I’ve used, this one was about as easy to install as a single click. It also ran smoother than any of the others I’ve tried, which was nice. That said, the output quality just isn’t good enough for me yet.
Has anyone else here used OmniVoice? I’d love to hear what kind of videos you tried and how the quality turned out. I ran it on a Mac, so I’m also curious to hear from people who’ve used it on other setups.
What I liked
- About as easy to install as one click (the easiest of any open-source tool I've tried)
- Ran more smoothly than any open-source tool I've tried so far
- Fast processing (about 3 minutes for a 22-second clip)
What I didn't
- The output just isn't good enough yet
- Cloning a voice into another language sometimes breaks into robot noise
Rating
Get the weekly AI dubbing digest
A weekly roundup of AI dubbing & news, in the language you pick. No spam, unsubscribe anytime.
Comments (0)
No comments yet — be the first.