Remove filler words: every um, uh and long pause.
Drop a podcast, a lecture, a voice-over or a video. AI finds the ums, uhs and dead air, shows each one in the transcript, and cuts the ones you leave ticked, with clean joins. It all runs in your browser; nothing is uploaded.
drop a recording or a video
MP3, WAV, M4A, FLAC and most other audio, or an MP4, MOV, MKV or WebM video. One file, up to two hours.
Struck through = will be cut. Click a word to keep or cut it; shift-click to hear it in context. ⏸ marks a long pause.
At a glance
| Finds | um, uh, ah, er, erm and hmm as Whisper writes them; held "mmm" and "uhhh" sounds it hears as a word, by their flat pitch (offered, not ticked); pauses over a second |
|---|---|
| You decide | Every find is a toggle in the transcript, with a listen-around preview |
| Cuts | Placed on the audio, not on the word times; joined with a 10 ms crossfade at a quiet sample; the pause left behind is at most a quarter of a second |
| Takes | Any audio the site reads, or an MP4, MOV, MKV or WebM video, up to two hours |
| Gives | WAV or FLAC at the source's own rate and depth, MP3 V0, M4A; a video as MP4 with the picture cut at the same places |
| Model | Whisper base with word timing, 135 MB the first time, then kept by the browser |
| Uploads your audio | No |
How it finds an "um"
The recording is transcribed in your browser by Whisper, the open speech model, in a version that also gives the time of every word. Whisper is trained to write tidy text and often leaves fillers out of a plain transcript, but with word timing on it writes most of them: in our test set, eight of nine ums and uhs came back as words, each starting within about a tenth of a second of where it really did. Every word on the filler list (um, uh, ah, er, erm, hmm and their stretched spellings) is marked and ticked.
Whisper sometimes hears a long "mmm" as a real word ("bum" was one in the tests). So a second pass looks at the short words that stand apart, between commas or pauses, and measures their pitch: a filler is one held note, while a real word moves. Those are marked as possible fillers and left unticked, because "so," and "well," at the start of a sentence can sound much the same, and cutting a real word is a worse mistake than leaving an um in. One case in nine was missed entirely: an "uh" Whisper folded into the word before it. Listen through the result before you publish it.
Why the cuts sound clean
Word times are only used to find a filler. The cut itself is placed on the audio: the page measures the recording's own background level, starts the cut where the filler's sound begins and ends it where the sound has died away, which matters because Whisper's word ends run early and an "um" trails off in a quiet hum. Each edge moves to the quietest sample within 5 ms, and the two sides are joined with a 10 ms equal-power crossfade, so there is no click. The silence either side of the filler is trimmed so that the pause left behind is at most a quarter of a second: taking out an "um" should not leave a hole where it was.
Long pauses are handled separately. A gap of more than a second between words is shortened to half a second, but only where the gap is really quiet, so a word Whisper missed is never cut out with it. Untick "Shorten pauses" to keep your timing exactly.
Podcasts, lectures and videos
For a podcast or an interview, export each speaker's track or the final mix and run it here before mastering. WAV and FLAC come out at the file's own sample rate and bit depth, so nothing is lost before you level it with the loudness normalizer. For a video, the picture is cut at the same places as the sound so lips stay in sync. That means the video is re-encoded (H.264 at high quality), unlike the site's other video tools, which copy the picture untouched. A jump cut where a filler was is normal in talking-head videos; in a screen recording it is usually invisible.
What it can't do
It removes sounds, not habits: "like", "you know" and "so" are real words to the model, and the page will not cut them, because whether one is filler depends on the sentence. Use the transcript to find them and the audio editor to cut them by hand. When two people talk over each other, one person's filler may sit under the other's word, and those cuts are best left unticked. Transcription runs on your computer, so a long file takes a while: roughly the length of the recording on a laptop without a fast GPU. Other languages work, but the filler list is English-led; Spanish "eh" is found, other languages' fillers may not be.
FAQ
How do I remove ums and uhs from audio for free?
Drop the recording above and press Find the fillers. Each um and uh is struck through in the transcript; click any you want to keep, then press Cut them and download. Nothing is uploaded and there is no watermark or account.
Does it work on video?
Yes. Drop an MP4, MOV, MKV or WebM. The sound is cut and the picture is cut at the same places, so lips stay in sync, and you get an MP4 back. The video is re-encoded to do that.
Will it cut real words by mistake?
It only ticks words on the filler list. Possible fillers it finds by sound are shown but left unticked, and long pauses are only shortened where the gap is quiet. Check the struck-through words before cutting; shift-click any of them to hear it.
Can it remove "like" and "you know"?
No. Those are real words and whether one is filler depends on the sentence, so they are left alone. Find them in the transcript and cut them in the audio editor.
Is my recording uploaded?
No. The speech model is downloaded to your browser and the recording is transcribed and cut there. The only download is the model, which is the same for everyone.
How long does it take?
The first run downloads a 135 MB model. After that, transcription takes roughly the length of the recording on a typical laptop, and less with a fast GPU; the cutting itself takes seconds.