opendub.ai
#persodub#ai-dubbing#open-source#voice-cloning

I Built a Dubbing App in Two Weeks and Shipped It

opendub · 2026-08-10 · 5 min read

My last post was July 14, so it has been almost a month. What I was doing in between: I moved from trying dubbing tools to building one.

I have used a lot of open-source dubbing tools. I wrote them up here, and I found bugs and sent fixes. Along the way it became clear what bothered me about all of them, and at some point I thought I might as well build my own. So I did. It is called PersoDub, a desktop dubbing app, and the first version went out on August 7.

The audio kept speeding up

There was one reason I built it.

When you dub a video, the translated line often runs longer than the original. Three seconds in Korean can take four in English. Most tools deal with that by speeding the audio up to squeeze it back into three seconds. Once or twice you don’t notice. Across a whole video you do, and it stops sounding like someone talking.

PersoDub doesn’t do that. Instead of touching the audio, it handles the length during translation. Each line is translated to fit the slot it came from, and when the video is put back together, any line that isn’t playing at exactly 1.000x gets written into the log. It isn’t a rule I try to remember. The code checks it.

Two weeks, minus sleep

The goal was simple. Not to build something perfect, but to get a usable first version out fast.

So for two weeks every hour went into it except the ones I slept. I started when I woke up and kept going until I went to bed. Two weeks later there was something to release.

Testing for quality was the hard part

Harder than building it was deciding when something was good enough.

Picking the model that generates the voices took a long time on its own. I would wire up a candidate, run the same lines through it, then listen to the cloned voice next to the original to judge how close it was. That is done by ear, so it never finishes in one pass. Same with the model that turns speech into text. I fed the same video to each one and compared which made fewer mistakes.

None of it was quick. When something didn’t work I stayed on it until it did, and every time I thought “one more try and I have it,” three or four hours had gone by.

I kept wandering off

The other hard part was order.

I would need to move on to the next stage, but a small bug would catch my eye and I would fix that first. Fixing it would show me another one. Big things and small things kept arriving together, and deciding what came first was harder than the building. Looking back, what made two weeks possible wasn’t speed. It was settling on “this has to work in this version, extra features can wait for the next release” and picking what shipped.

What came out

Original (English)

Dubbed with PersoDub (Korean)

The same clip, dubbed from English into Korean. The speaker's own voice is cloned.

You drop in a video, and it separates the background from the voices, transcribes the speech, works out who spoke when, translates each line, clones each speaker’s voice and lays it back over the original. Music and effects stay where they are, only the speech changes. With the default settings all of it runs inside my MacBook Air M4. No account, no API key, and no internet except for the one-time model download.

It doesn’t support every OS yet

Right now it only runs on Apple Silicon Macs. Rather than half-supporting every OS, I wanted it to work properly on one. Other systems will come later.

One line to install

Installing this kind of thing has always been painful. You download one piece, and it tells you to download another. I put the whole thing into one line.

curl -fsSL https://raw.githubusercontent.com/stronghamjji/PersoDub/HEAD/install.sh | bash

The code and the install steps are on GitHub. It is a first version, so there will be bugs. Please open an issue if you run into one. I’ll keep posting about how it’s going here too.

comments

Comments (0)

No comments yet — be the first.