Transcribe audio to text, free and private.
Drop a recording or a video and get a transcript, an SRT or a VTT subtitle file: MP3 to text, voice memo to text, video to text, with each speaker labelled if you want it. The speech model is OpenAI's Whisper, and it runs inside this browser tab — your audio is never uploaded, so an interview, a medical note or an unreleased song stays on your computer. 99 languages, or translate any of them into English.
drop audio or video
MP3, M4A, WAV, voice memos, MP4, MOV, WebM. One file at a time, up to two hours.
At a glance
| Model | OpenAI Whisper (tiny, base or small), ONNX, through transformers.js |
|---|---|
| Languages | 99 for transcription; translation goes into English only |
| Outputs | Plain text (editable before you save), SRT and WebVTT subtitles |
| Model download | Once, then cached: about 90 MB (tiny), 135 MB (base), 285 MB (small) |
| Speed | Whisper base measured at about a sixth of the audio's length on a laptop GPU and under a third on its CPU: a ten-minute interview in two to three minutes |
| Timestamps | Per phrase, roughly every few seconds — not per word |
| Length limit | About two hours per file |
| Uploads your file | No. The only download is the model, which is the same for everyone and holds none of your audio |
Unlimited transcription, because nothing is uploaded
The file is decoded in this tab, folded down to one channel and resampled to 16 kHz, which is the only rate Whisper understands. That happens with the same windowed-sinc resampler the rest of AudioSaw uses rather than the browser's built-in conversion, so there is no aliasing for the model to mistake for sibilance. The audio is then handed to a background worker that runs the network in 30-second windows with five seconds of overlap on each side, and stitches the windows back together where the overlaps agree.
None of that leaves your machine. Most free transcription sites upload the recording and run Whisper on their own server, which is why they want an account and cap you at a few minutes: someone is paying for the GPU. Here your GPU, or failing that your CPU, does the work, so there is no account, no minute cap and no copy of your audio sitting on a stranger's disk. That matters most for exactly the recordings people most want transcribed — interviews with sources, therapy and medical dictation, legal calls, lectures you were not supposed to record, a draft of a song.
Which Whisper model to pick
Base is the default because it is the best trade for most recordings: clear speech in English and the major European languages comes back with very few mistakes, and the download is a one-off. Tiny is the one to use on a phone or an older laptop, or when you only need to search a recording rather than quote from it; it is several times faster and noticeably sloppier with names, numbers and accents. Small is the most accurate option here and the one to choose for a heavy accent, a noisy room, or any language other than the big European ones — Hindi, Persian, Tamil and Vietnamese all improve markedly from base to small. It is a large download and slow on a CPU.
The page probes your browser before you start and says whether it can use the GPU. With WebGPU (current Chrome and Edge on a desktop) Whisper runs several times faster than on the CPU; the words are the same either way. Without it, the CPU path splits the work across your cores, which is why the tab needs to stay open and why a laptop fan may spin up.
Video to text: SRT or VTT subtitles
Both files hold the same cues — a start time, an end time and a line or two of text. SRT is the one nearly every editor and platform imports: YouTube Studio, Premiere Pro, DaVinci Resolve, Final Cut (via its caption import), CapCut and VLC all take it. VTT is the web's own format and what an HTML5 <track> element needs. Lines are wrapped at 42 characters, the broadcast convention, and every cue is kept on screen for at least a third of a second so nothing flashes past unread.
Whisper times phrases, not words, so a cue covers a natural chunk of speech of a few seconds. That is the right granularity for subtitles. It is not the right tool for karaoke-style word highlighting, and the page does not pretend otherwise.
What it gets wrong
Proper nouns are the commonest error: a product, a surname or a place it has not heard spelled is written the way it sounds. That is why the transcript is an editable box — fix the names before you download the text. Subtitle files keep the original wording so their timing stays aligned.
Transcribe an interview with speaker labels
Whisper itself does not tell speakers apart; tick "label who is speaking" and a second step does. It is pyannote's method, run in your browser: a small segmentation model (pyannote segmentation-3.0) marks where speech is and when the voice changes, a speaker-embedding model (WeSpeaker ResNet34, CC-BY-4.0) turns each stretch of speech into a voiceprint, and the voiceprints are grouped into people. Each line of the transcript then gets "Speaker 1:", "Speaker 2:" and so on, and you can type real names in once. Subtitles get the name at each change of speaker. On test conversations it labelled over 99% of the speech correctly with two people and 92% with three, and counted them right when left to count; tell it the number of people if you know it. Short interjections ("yes", "right") are the usual mistakes, and two people with very similar voices may be merged. Crosstalk, where two people talk at once, tends to keep one voice and lose the other.
What else it gets wrong
On long stretches of music or silence it can invent words — usually a repeated phrase or a polite "Thank you." It was trained largely on subtitled video, and that is what subtitles say when nobody is talking. Trimming a long music intro with the cutter, or tightening the gaps with auto-cut silence, removes most of it. Background hiss and hum hurt accuracy more than people expect; a pass through noise reduction first is often the single biggest improvement.
"Translate into English" is Whisper's own translation, done in the same pass. It is good enough to understand a foreign-language recording and to make rough English subtitles; it is not a translator you would publish from unchecked, and it only goes into English.
Compared with the alternatives
Otter, Rev, Descript and the transcription built into meeting apps are convenient and label speakers too, but they all work by uploading your audio and most charge by the minute beyond a free allowance. Running Whisper yourself in Python, or with a desktop app such as MacWhisper, gives you the larger models and is the better choice for hours of material every week. This page sits between the two: the same model family as the desktop route, with nothing to install, and the privacy of doing it on your own machine. If you do want to run Whisper elsewhere, audio for Whisper produces the 16 kHz mono WAV it is fed internally.
FAQ
Is my audio uploaded anywhere?
No. The speech model is downloaded to your browser and runs there; your recording never leaves the tab. The only network request is for the model files, which are identical for everyone and contain nothing of yours.
Is this really OpenAI's Whisper?
Yes — the open Whisper models OpenAI released, converted to run in a browser with ONNX Runtime. You choose between the tiny, base and small sizes; the larger desktop models are too big for a browser tab.
Can I transcribe an MP3 to text for free, with no time limit?
Yes. There is no account and no minute allowance, because your own computer does the work. Files of up to about two hours go in one piece; split anything longer with the split audio tool first. MP3, M4A voice memos, WAV, FLAC, OGG, WhatsApp voice notes and video files all work.
How long does it take?
Whisper base measured at about a sixth of the audio's length on a laptop GPU and under a third on its CPU, so a ten-minute recording takes two to three minutes. The first run also downloads the model, once. The page estimates the time for your file before you start.
Can it make subtitles for a video?
Yes. Drop the MP4, MOV or WebM directly — the audio is pulled out in the browser — and download an SRT for YouTube, Premiere, Resolve or CapCut, or a VTT for the web.
Does it label who is speaking?
Yes, if you tick "label who is speaking". A speaker model (pyannote segmentation with WeSpeaker voiceprints) runs in your browser after the transcription and marks each line "Speaker 1", "Speaker 2" and so on, which you can rename. On test conversations it labelled over 99% of the speech correctly with two people and 92% with three.
Which languages does it support?
Whisper covers 99 languages and detects which one is spoken. Accuracy is best in English and the major European languages; for others, pick the language yourself and use the small model. "Translate into English" turns any of them into English text.
Will it work on my phone?
The tiny model will, slowly. Phones have little memory for the larger models and long files, so for anything over a few minutes a laptop or desktop is the better choice.