Speech to text: speak, and it types.
Online dictation and voice typing: press the button, talk, and your words appear as text, with punctuation, after each phrase. Voice to text in 99 languages, free, with no sign-up and no time limit. The speech recognition is OpenAI's Whisper running inside this browser tab, so nothing you say is sent to a server. That is the difference from the dictation built into Chrome, Google Docs and most phones, which stream your voice to the cloud.
The first time, the speech model downloads (135 MB for the recommended one) and is kept for next time.
At a glance
| Recognition | OpenAI Whisper (tiny, base or small) through transformers.js, running in the tab |
|---|---|
| Languages | 99 to write in the language spoken, or any of them translated into English |
| How text arrives | A phrase at a time: after each pause of about 0.7 seconds, with punctuation and capitals |
| Uploads your voice | No. Audio goes from the microphone to the model in this tab and is then discarded |
| Keeps | Only the text, in this browser, so a reload does not lose it |
How voice typing works here
The microphone is read in small blocks and measured every 30 milliseconds against the room's own background level, which the page keeps learning while you're quiet. Anything 10 dB above that background is speech. When you pause for about seven-tenths of a second, the phrase you just said is handed to Whisper, the same speech model the audio to text page uses for recordings, and the text it returns is added to the box. The 300 milliseconds before you started speaking are included, so the soft start of a word like "first" or "hello" isn't clipped. A phrase that runs past 25 seconds is cut there, inside the 30-second window Whisper listens through.
Whisper is given whole phrases rather than fixed slices of time, which is why the text arrives with sensible punctuation: it sees the sentence it's punctuating. That is also why it lags by a phrase rather than showing each word as you say it. Streaming word-by-word recognisers exist, but they run in the cloud.
Getting accurate text from talk to text
Speak in phrases and pause at the end of each sentence; the pause is what sends the phrase. A headset or a laptop's built-in microphone in a quiet room works well; a phone on a desk across the room does not. Say punctuation only if you want the words: Whisper adds commas and full stops from your intonation, and "comma" is written as the word. Pick your language rather than leaving it on automatic for short phrases, because a two-word phrase is not much to identify a language from.
Base is the right model for most people. Tiny keeps up better on an older laptop or a phone, with more mistakes on names and numbers. Small is noticeably better with accents and for languages other than English, but on a computer without a GPU it may fall several phrases behind; it catches up when you stop.
Private dictation: your voice is never uploaded
Most dictation is a cloud service. The microphone button in Google Docs, Chrome's speech API, Windows voice typing in its default mode and most phone keyboards send the audio to the vendor's servers to be recognised. That is fine for a shopping list. It is less fine for a patient note, a legal memo, a diary or a draft that isn't public yet. Here the model is downloaded once and runs on your machine; the audio is discarded as soon as it has been transcribed, and only the text is kept, in your browser, until you clear it.
What it can't do
It does not type into other apps: the text appears in this box, and you copy it from there or download it as a .txt file. It does not take voice commands such as "new paragraph" or "delete that" — a long pause starts a new paragraph, and you edit in the box like any text. It does not label different speakers. For a recording you already have, use audio to text, which also writes SRT and VTT subtitles. To hear text read aloud instead, there is text to speech.
FAQ
Is my voice sent anywhere?
No. The speech model runs in your browser; the audio never leaves the tab and is thrown away once it has been turned into text. Only the model download touches the network, once.
Why does the text appear after I pause rather than word by word?
Each phrase is transcribed as a whole when you pause, which is what gives Whisper enough context to punctuate it. Word-by-word recognition needs a different kind of model that today only runs on servers.
Is this speech to text free, with no sign-up and no time limit?
Yes. There is no account and no minute allowance: the speech model runs on your own computer, so dictating for an hour costs the same as for a minute. The one download is the model, once.
Can I transcribe a recording instead of talking live?
Yes, on the audio to text page: drop an MP3, a voice memo or a video and it writes a transcript or SRT subtitles, and can label who is speaking. This page is for live voice typing from the microphone.
Which languages can I dictate in?
Whisper covers 99 languages. Choose yours from the list for the best results, or choose "An English translation" to speak another language and get English text.
Is the text saved if I close the tab?
The text is kept in this browser and comes back when you return to the page, until you clear it. Download it as a .txt file to keep it anywhere else.
Does it work on a phone?
On recent phones, with the tiny or base model. Text arrives a little slower than on a laptop.