AI video dubbing and translation, free.
An AI video translator with no account and no upload: drop a video in any language and get it back with a voice-over in English, Spanish, French, Italian, Brazilian Portuguese or Hindi, timed to the original speech. Whisper transcribes the soundtrack and translates it, you check the lines, and an AI voice reads them over the original with the music and effects kept underneath. All of it runs in this tab: the video is never uploaded, so this works for footage you can't share with a cloud service.
drop a video
MP4, MOV, WebM, or an audio file. Up to about 600 MB.
The English lines, with their times — fix anything before the voice reads it
At a glance
| Speech recognition | OpenAI Whisper (base or small), translating from 99 languages into English |
|---|---|
| Translation out of English | Helsinki-NLP OPUS-MT, one model per language (about 110–130 MB, downloaded once): Spanish, French, Italian, Brazilian Portuguese, Hindi |
| Speakers | One narrator, or a different voice for each person (speakers found automatically) |
| Voice | Kokoro-82M: 28 English voices, plus 13 Spanish, French, Italian, Brazilian Portuguese and Hindi voices; or, for English, each speaker's own voice cloned from the video (Chatterbox Turbo, WebGPU) |
| Timing | Each line starts where the original phrase started; a line that runs long is read up to 1.3× faster, and anything still over is flagged for you to shorten |
| Original audio | Turned down 18 dB under each line, kept at full level, or replaced |
| Video | Copied untouched, not re-encoded: same picture, same quality |
| Outputs | Dubbed MP4, the new soundtrack as MP3, subtitles in the dub's language as SRT |
| Uploads your video | No — everything runs in this tab |
How to dub a video into another language for free
The soundtrack is decoded in the browser and handed to Whisper, which returns the speech as timed phrases, translated into English if you asked for that. Fragments are joined into natural lines, but never across a pause or a full stop. Whisper only translates into English, so a dub into Spanish, French, Italian, Portuguese or Hindi takes one more step: the English lines go through OPUS-MT, the University of Helsinki's open translation models, one sentence at a time (given two sentences at once, the model sometimes drops one). That runs on the processor at about a second a line, after a one-time download of the language's model. You then see the lines in a list with their times, and in a translated dub each one has the English it came from underneath. This is the moment to fix a name or soften a clumsy translation. Machine translation of speech is good enough to follow a video, not good enough to publish unread.
Then each line is read by Kokoro, an open neural voice, and placed where the original phrase began. A translation often runs longer than the language it came from, so a line that would overrun the start of the next one is read again slightly faster, up to 1.3×, which is about as fast as a voice stays natural. Past that, the line is allowed to run on and is marked in the list, so you can shorten its wording and dub again. The original soundtrack is turned down by 18 dB under each line and left alone in between, so music, effects and the presence of the original speakers stay; that is how most documentary voice-over is mixed. Finally the new soundtrack is muxed onto the video with the picture copied bit for bit, so nothing is re-compressed.
Translating a video to English or Spanish: good results
Clear speech with modest background works best; a single presenter, an interview, a lecture, a cooking or how-to video. For a translation, choose Whisper small: its translations are noticeably better than base's, at the cost of a bigger download and more time. Pick the spoken language yourself if the video starts with music. Choose a voice that suits the speaker; Michael and Fenrir are steady male narrators, Heart and Bella the most natural female voices. If the original speech is still too present under the new voice, choose "Replace it", or separate the voice from the music first with the AI vocal remover and dub the instrumental.
What it can't do
The target languages are the six that Kokoro speaks: English, Spanish, French, Italian, Brazilian Portuguese and Hindi. German is out because Kokoro has no German voice; its Japanese and Chinese voices need a pronunciation engine that has no browser version yet. Going into a language other than English, the speech is translated twice (Whisper into English, then OPUS-MT out of it), so a subtlety lost in the first step stays lost; for a Spanish video you want in French, check the lines against the original. On a four-line how-to script, the Spanish, French, Italian and Portuguese translations were all right; the Hindi one got the meaning of one line wrong, so a Hindi dub especially needs a reader who knows the language. French has a single voice, so in a French dub every speaker shares it. Re-voicing without translating works in the same six languages: a Spanish video can be given a new Spanish voice (measured: a Spanish clip re-voiced in Spanish transcribes back with every key word). With "a different voice for each speaker" ticked, a speaker model (pyannote with WeSpeaker voiceprints, the same one behind the speaker labels on audio to text) works out who says each line, and every person gets their own voice, chosen per speaker; by default they alternate a man's and a woman's voice so two people stay easy to tell apart. Without it, one narrator reads everything. With "use each speaker's own voice" ticked, an English dub is read in voices cloned from the speakers themselves: about five seconds of each person's own lines is the reference, and a line too long for its slot is tightened at the same pitch. On a test clip of a woman's voice, the dub came out in her register (a median of 171 Hz, against about 120 for the default male voice) and Whisper read it back nearly word for word. It needs WebGPU, downloads the 560 MB cloning model once, takes about four and a half times the speech's length, and must only be used with the speakers' permission; music under the voice muddies the clone, so separate it first. Without it, the voice does not match the original speaker's voice. Either way the dub is not lip-synced: this is voice-over dubbing, as on news and documentaries, not the studio kind where actors match mouth movements. Cloud services such as ElevenLabs and Rask do voice cloning and lip-sync by uploading your video to their servers and charging by the minute. Long videos take a while: transcription runs at a fraction of the video's length on a GPU, and the voice at about the length of the speech. Trim to the part you need with the cutter or your editor first.
If you only need English subtitles, audio to text writes them directly. The SRT from this page is in the dub's language and loads into YouTube, Premiere or CapCut alongside the dub.
FAQ
Is my video uploaded?
No. The video is read, transcribed, voiced and re-assembled in this browser tab. The only downloads are the speech and voice models, which are the same for everyone.
Is there a free AI video translator with no sign-up?
This is one. There is no account, no credits and no watermark, because the speech recognition, translation and voice all run in your browser; the only downloads are the models, once. Videos up to about 600 MB work, and longer ones take a while on a CPU.
Can I get subtitles in another language from a video?
Yes. Every dub also saves an SRT subtitle file in the dub's language, which YouTube, Premiere and CapCut import. For English subtitles alone, without a voice-over, audio to text writes them directly.
Which languages can it dub from?
Whisper understands 99 languages and translates them into English. Accuracy is best for widely spoken languages; for others, pick the language and use the small model.
Can it dub into Spanish, French or other languages?
Yes: Spanish, French, Italian, Brazilian Portuguese and Hindi, as well as English. Whisper translates the speech into English, and an open translation model (OPUS-MT, about 110–130 MB, downloaded once per language) takes it on from there. Other languages are not offered because the voice model does not speak them.
Does it re-encode my video?
No. The picture is copied exactly; only the soundtrack is new. The file comes back as MP4.
Can each person in the video get their own voice?
Yes. Tick "a different voice for each speaker": the page works out who says each line and offers a voice for each person. On a two-person test it found both speakers and gave each line the right voice.
Can the dub keep the original speaker's voice?
For an English dub on a computer with WebGPU, yes: tick "use each speaker's own voice" and each person's lines are read in a voice cloned from their own speech in the video, with the same open model as the voice cloning page. Only do this with the speakers' permission. Otherwise it uses one of the 41 built-in voices.
Can I fix the translation before it is spoken?
Yes. After transcription every line is shown in an editable box with its time; the voice reads whatever is in the boxes when you press Make the dub.