Free text to speech, with MP3 download.
Type or paste text and hear it read aloud in a natural AI voice, then download it as an MP3 or WAV: text to audio, with no sign-up and no daily character allowance to run out of. The voice generator is Kokoro, an open neural text-to-speech model, and it runs inside this tab, so the text is never sent anywhere. 41 voices speak American and British English, Spanish, French, Italian, Brazilian Portuguese and Hindi. Type, paste, or open a PDF, Word document, text file or EPUB and have it read aloud.
It plays each sentence as soon as it is ready. The first run downloads the voice model and keeps it; after that, speech starts within seconds. Every file is labelled as synthetic speech in its own tags.
At a glance
| Model | Kokoro-82M (Apache-2.0), run with ONNX Runtime Web |
|---|---|
| Voices | 41: 20 American and 8 British English, 3 Spanish, 1 French, 2 Italian, 3 Brazilian Portuguese and 4 Hindi, plus any blend of up to three |
| Speed | Measured on a laptop (Apple GPU, 8 cores): about half the speech's own length on the GPU, about one and a half times it on the CPU. A busy machine takes longer |
| Model download | Once, then cached: 326 MB for the fast GPU model, or 92 MB for the smaller one that runs on the CPU |
| Output | 24 kHz mono — MP3, WAV or M4A |
| Limits | 20,000 characters per run here; longer texts go to the audiobook maker |
| Uploads your text | No. The only network traffic is the model and voice files, which are the same for everyone |
Why it is free and unlimited: it runs on your computer
Your text is cleaned up first: numbers, prices, times and titles are written out the way they are said, so "$4.50" becomes "four dollars and fifty cents" and "Dr. Lee" becomes "Doctor Lee". It is then split into sentence-sized pieces, because the model is at its best on a sentence or two at a time and starts to rush through very long inputs. Each piece goes through espeak-ng, a long-standing open pronunciation engine, which turns spelling into phonemes, and then through Kokoro, which turns the phonemes into a waveform in the chosen voice.
All of that happens in a background worker in this tab. Most free text-to-speech sites send your text to a cloud API and pay per character for the result, which is why they cap you, ask for an email address, or put the download behind a plan. Here your own GPU, or your CPU if there is no usable GPU, does the work. A contract, a medical letter, an unpublished chapter or a script under NDA stays on your machine.
Many files at once: one MP3 per line
"One file per line" turns every line in the box into its own audio file and hands you a zip: the prompts for a phone menu, the voice lines of a game, the slides of a course, a list of vocabulary words. Write a line as name | text to choose its file name, or paste two columns from a spreadsheet (file name, then text), and the files come out named that way; a line with no name gets its number and first words. Up to 200 lines go in one batch, with the voice, speed and format you have chosen, and every file is tagged as synthetic speech like a single download.
Read a PDF or Word document aloud
"Open a file" puts the text of a PDF, a Word document (.docx), a plain text file or an EPUB into the box, where you can trim it before pressing Speak. A PDF is rebuilt into paragraphs the way the audiobook maker does it: lines joined, words hyphenated across a line break put back together, and the running header and page number that repeat on every page left out, so the voice does not read "Page 12" between sentences. Chapter and section titles stay as their own lines, which gives a natural pause before each one. A scanned PDF, which holds pictures of pages rather than text, needs OCR first, and the page says so. The box holds 20,000 characters, roughly twenty minutes of speech; a longer file loads its first part, and the audiobook maker reads the whole thing into one file with chapters.
Choosing a natural AI voice
Each voice carries a grade from the model's authors, shown in the list. It reflects how much clean training audio that voice had, and it is a fair guide: Heart and Bella (both American, female) are the most natural voices in the set, Emma and Nicole are close behind, and the voices graded D have audible roughness on some words. Among the male voices, Michael, Fenrir and Puck are the steadiest. "Hear this voice" plays a short sample before you commit to a long text.
Mixing blends voices rather than switching between them. Each Kokoro voice is a table of numbers that steers the model, and a mix is a weighted average of those tables, so Bella at two parts and Michael at one gives a new, consistent voice somewhere between the two. It is a way to get a voice nobody else's video is using. It is not voice design from a written description, and it cannot reproduce a particular person. The accent always follows the first voice in the mix.
Making it sound natural, then saving the MP3
Punctuation is direction. Commas give short breaths, full stops give longer ones, a question mark lifts the end of a line, and a blank line between paragraphs gives a longer pause. If a sentence comes out flat, splitting it in two usually fixes it. Spell out anything with an unusual pronunciation the way it sounds. A brand or surname the model misreads can be respelled phonetically, and the respelling never shows up anywhere but the audio. Speed between 0.9× and 1.1× sounds the most natural; outside that range the voice stays in tune but its rhythm starts to sound processed.
For a voice-over, generate a paragraph at a time so you can redo one line without regenerating everything, then line the pieces up in the audio editor against your music. For a podcast intro or a video narration, finish with the loudness normalizer so the voice sits at the level the platform expects.
Hindi, Spanish, French, Italian and Portuguese text to speech
Pick a voice from the language's group in the list and type in that language: Hindi in Devanagari, Spanish with its accents. There are four Hindi voices (Alpha and Beta, female; Omega and Psi, male), three Spanish (Dora, Alex, Santa), three Brazilian Portuguese with the same names, two Italian (Sara and Nicola) and one French (Siwis). The first time you choose one of them, the page downloads a fuller pronunciation engine (espeak-ng, 18 MB) that knows how those languages are spelled. Read back by Whisper, the Spanish and Portuguese test sentences came out word for word, French and Italian missed only numbers written as digits (write them out to be safe), and the Hindi was correct. The same 20,000 characters per run and the same MP3 download apply, with no word limit across runs. To have a Spanish or Hindi version of a video rather than of a text, video dubbing translates and voices it in one go.
What it can't do
There is no cloning on this page: the voices are Kokoro's own. To make speech in a particular person's voice, with their permission, use voice cloning, which needs a GPU and a much bigger model. Beyond English, Kokoro speaks the five languages below, with fewer voices each than English, and Japanese and Chinese need a pronunciation engine that does not run in a browser yet. Emotional delivery is limited: Kokoro reads in a calm, even narrator's style and does not do shouting, whispering or crying on request.
Generated files say so in their own metadata (an ID3 comment in an MP3, an INFO chunk in a WAV) so that a file found later can be traced. That label is easy to remove and is not a watermark in the audio itself. Please don't use synthetic speech to impersonate real people.
Compared with the alternatives
Your operating system already has a reader: macOS's Spoken Content, Windows Narrator and Chrome's "read aloud" all speak text, but their voices are older-generation, and only macOS's command-line say saves a file. Cloud services (Google Cloud TTS, Amazon Polly, Azure, ElevenLabs) have more voices and languages and charge per character beyond a free tier, with an API key or account either way. Kokoro sits in between: a modern neural voice rated close to much larger models in public listening comparisons, small enough to run on an ordinary laptop, and licensed so that the audio you make is yours to use.
FAQ
Is it really free and unlimited, with no sign-up?
Yes. There is no account, no login and no monthly character quota: the work happens on your own computer, so there is nothing per character to pay for. The one limit is 20,000 characters per run on this page, to keep the tab responsive; run it again as often as you like, and the audiobook maker handles whole books chapter by chapter.
Can I download the speech as an MP3?
Yes. Choose MP3 (192 kbps), WAV or M4A and press Download once it has finished speaking. The file is saved straight from the tab, with no watermark in the audio and no email or plan needed to get it.
Is my text uploaded anywhere?
No. The voice model is downloaded to your browser and runs there. Your text never leaves the tab; the only requests are for the model and voice files, which are the same for everyone.
Can I use the audio commercially?
Kokoro is released under the Apache-2.0 licence, which allows commercial use, and nothing on this page adds restrictions. You are responsible for what the speech says and for not passing it off as a real person.
Why is the first run slow?
It downloads the voice model: 326 MB for the fast GPU version, or 92 MB if you choose the smaller download or have no usable GPU. It is cached afterwards, so the next text starts speaking within a few seconds.
Can it clone my voice?
Not on this page, which uses Kokoro's built-in voices (you can blend them with "Mix voices"). The voice cloning page does it from five seconds of your recording, with a bigger model that needs a GPU, and only for your own voice or one you have permission to use.
Can I make many MP3 files at once from a list?
Yes. Put one line per file in the box, optionally as "name | text", and press "One file per line". Each line becomes its own MP3 (or WAV or M4A), and they come as one zip. Up to 200 lines per batch.
Can it read a PDF or a Word document aloud?
Yes. Press "Open a file" and choose a PDF, a .docx, a .txt or an EPUB. The text goes into the box with headers and page numbers removed, and you can listen or download it as an MP3. For a whole book with chapters, use the text to audiobook page.
Is there Hindi or Spanish text to speech?
Yes. American and British English have 28 voices, and Spanish, French, Italian, Brazilian Portuguese and Hindi 13 more between them, including two female and two male Hindi voices. Japanese and Chinese need a pronunciation engine that does not yet run in a browser.
What is Kokoro TTS, and can I use it online?
Kokoro-82M is an open-weight text-to-speech model, released under Apache-2.0 and small enough (82 million parameters) to run on a laptop and rated close to much larger models in public listening comparisons. This page is Kokoro running in the browser: all of its English, Spanish, French, Italian, Portuguese and Hindi voices, with nothing to install, no Python and no GPU server.
Does it work on a phone?
On recent phones, yes, but slowly, because without a desktop-class GPU the model runs on the processor. Short texts are fine; for long ones a laptop is much quicker.