Free voice cloning, in your browser.
Record a few seconds of a voice, or drop an audio or video file of it, type some text, and hear it spoken in that voice: an AI voice clone with no account, no credits and no upload. The cloning model is Chatterbox Turbo, an open model from Resemble AI, and it runs on your own graphics card inside this tab, so the voice sample is never uploaded to anyone. Use it for your own voice, or a voice whose owner has said yes.
1. The voice to clone
Record yourself reading this aloud, in a quiet room, at your normal pace:
The north wind and the sun were disputing which was the stronger, when a traveller came along wrapped in a warm cloak. They agreed that the one who first made the traveller take his cloak off should be considered stronger than the other.
drop a voice clip
WAV, MP3, M4A, a voice memo. The first five seconds of speech are used.
2. What it should say
At a glance
| Model | Chatterbox Turbo (Resemble AI, MIT licence, 350M parameters), 4-bit weights, run with ONNX Runtime Web |
|---|---|
| Reference | One person speaking; the first five seconds of clean speech are used (three at the least) |
| Language | English |
| Needs | WebGPU: current Chrome or Edge on a desktop or laptop. About 560 MB downloaded once and cached |
| Speed | Measured on a laptop GPU (Apple, 8 cores): about four and a half times the length of the speech, so a 4-second line takes about 17 seconds. The first one also loads the model |
| Output | 24 kHz mono MP3 or WAV, tagged as a cloned synthetic voice |
| Uploads your voice | No. The sample and the text stay in this tab |
How zero-shot AI voice cloning works
Zero-shot cloning means the model has never been trained on your voice; it hears a sample once and imitates it. Chatterbox does this in three steps. A speech encoder listens to your reference clip and turns it into two things: a compact description of the voice (its pitch range, timbre and the way it sounds in its room), and the clip itself as a string of speech tokens. A language model then reads your text and writes new speech tokens, one for every 40 milliseconds of audio, continuing in the style of the reference. Finally a decoder turns those tokens back into sound in the reference voice.
All three run here, on your graphics card, through WebGPU. The model is downloaded from Hugging Face the first time, about 560 MB, and kept by the browser, so the second visit starts in seconds. Your recording is resampled to 24 kHz, trimmed of silence at the ends and cut to its first five seconds of speech; it is never sent anywhere, and closing the tab discards it.
How much audio do you need to clone a voice?
Five seconds of one person speaking clearly is enough here, and three is the minimum; a longer sample does not improve the likeness, for a reason explained below. What the reference sounds like matters far more than its length. Record in a quiet, soft room: a bedroom or a car beats a kitchen, because hard surfaces add echo, and the model copies the room along with the voice. Hold the phone or microphone a hand's width from your mouth and speak as you want the result to sound. If you read flatly, the clone reads flatly. Only the first five seconds of speech are used: the decoder re-reads the whole reference for every fraction of a second it renders, so a longer one makes every sentence slower without helping the likeness much. Record a little more than that and start speaking straight away. A clip with music, a second voice or heavy noise underneath gives a muddy clone; clean it first with noise reduction or pull the voice out of a song with the AI vocal remover.
Write the text the way it should be spoken, with punctuation, and keep sentences to a reasonable length; each one is generated separately and played as it arrives. Chatterbox Turbo also understands a few non-speech tags inline: [laugh], [chuckle], [cough] and [sigh].
Consent and honesty
A cloned voice can be used to deceive: fake phone calls to relatives, messages put in someone's mouth, attempts to get past a bank's voice check. Do not clone anyone without their permission. The page will not start until you confirm that the voice is yours or that you have permission, and every file it writes carries a tag saying it is a cloned, synthetic voice. That tag is metadata and easy to strip, so it is a label, not protection. The responsibility is yours. In many places, using someone's voice without consent is unlawful as well as wrong.
What it can't do
It needs WebGPU. On a phone, or in Safari or Firefox without it, the page says so instead of crawling for minutes on the processor. It speaks English only. The likeness is good, not perfect: people who know the voice well will usually hear the difference, especially in emotion and rhythm. For a neutral narrator voice with no cloning, the text to speech page is faster and needs no GPU. Commercial services such as ElevenLabs offer more languages and fine-tuned "professional" clones, but they work by uploading your voice to their servers.
FAQ
Is my voice uploaded?
No. The recording and the text stay in this browser tab, and the cloning model runs on your own graphics card. The only download is the model itself, which is the same for everyone.
How long a recording does it need?
Five seconds of one person speaking clearly. It uses the first five seconds of speech, and needs at least three.
Can I clone a voice from an audio file or a video?
Yes. Instead of recording, drop an MP3, WAV, M4A, voice note or video file of the person speaking alone. The first five seconds of speech are used, so trim the file to a clean stretch first if it starts with music or another voice.
Is this an open-source alternative to ElevenLabs?
For cloning in English, yes: Chatterbox Turbo is MIT-licensed, and here it runs on your own GPU with no account or credits. ElevenLabs has more languages, emotion control and trained "professional" clones, and uploads your voice to do it. AudioSaw's other voice tools cover the rest of that ground in the browser: text to speech, video dubbing, transcription and a voice changer.
What is Chatterbox Turbo, and can I try it online?
Chatterbox Turbo is Resemble AI's open voice-cloning model (MIT licence, 350 million parameters). This page runs it in the browser with 4-bit weights on WebGPU, so you can try it without Python, a GPU server or a Hugging Face account.
Why does it need WebGPU?
The model writes speech one small token at a time, 25 per second of audio, and each token is a pass through a 350-million-parameter network. A graphics card does that quickly; a processor alone would take minutes per sentence.
Can I clone a celebrity or someone else's voice?
Only with their permission. The page asks you to confirm it before it starts, and every file is tagged as a cloned voice.
Which languages does it speak?
English. Resemble AI's multilingual Chatterbox covers 23 languages but is too large to run comfortably in a browser yet.
Is it free?
Yes: no account, no credits, no limit. Your own computer does the work.