Enhance speech: one click to a cleaner voice.
Drop a voice recording, a podcast take, a lecture or a phone video and get it back sounding like it was recorded on purpose: background noise taken out by an AI denoiser, low rumble filtered, the voice made clearer, quiet and loud sentences evened out, and the whole thing brought to podcast or YouTube loudness without clipping. It runs in your browser, so nothing is uploaded, there is no sign-up and no minute limit.
drop a voice recording or video
MP3, WAV, M4A voice memos, FLAC, or an MP4, MOV or WebM video. Speech works best.
| Loudness before | — |
|---|---|
| Loudness after | — |
| Compression | — |
| True peak after | — |
| Background noise | — |
At a glance
| Noise | RNNoise, a small neural denoiser from Xiph, removes noise behind a voice, steady or not |
|---|---|
| Tone | 80 Hz high-pass (rumble, desk thumps), -2.5 dB at 250 Hz (boxiness), +3 dB at 3.5 kHz (presence) |
| Level | A gentle 3:1 compressor that brings quiet sentences up towards loud ones, then loudness to -16 or -14 LUFS (ITU-R BS.1770) with a -1 dBTP true-peak ceiling |
| Video | MP4, MOV or WebM in; the same video back with the enhanced sound, the picture untouched |
| Uploads your file | No — everything runs in this tab |
What it does to your recording, step by step
Most rough recordings have the same four problems, and this fixes them in the order a sound engineer would. First, noise: an AI denoiser (RNNoise, trained by Xiph on thousands of hours of speech and noise) keeps the voice and pulls down everything else, including the irregular noise a classic noise gate cannot touch, such as traffic, keyboards and a fan that changes speed. Second, rumble: a high-pass filter at 80 Hz removes the low thuds of a desk, a car or a handled phone, which eat loudness without adding anything you can hear as voice. Third, tone: a small cut at 250 Hz takes out the boxy sound of a small room, and a lift at 3.5 kHz brings forward the consonants that make speech intelligible on phone speakers and earbuds. Fourth, level: a gentle compressor narrows the gap between the loud and quiet moments, then the whole file is brought to a standard loudness, with a true-peak limiter so it never clips.
Measured on a test recording with a speaker who drops about 12 dB for one sentence, noise bursts and clicks under the voice, and a 40 Hz rumble: the result landed at -16.1 LUFS with its true peak at -1.0 dBTP; the quiet sentence came up, so the gap between it and the loud ones went from 11.5 dB to 6.8 dB; the pauses came down 24 dB relative to the voice; and the rumble all but disappeared. The voice stays where it was to a third of a millisecond and the file is exactly as long as before, so a video stays in sync. A 13-second clip took about 8 seconds on a laptop.
Enhance speech vs Adobe Podcast Enhance
Adobe's Enhance Speech is a generative model: it listens to your recording and re-synthesises the voice as if it were recorded in a studio. That can sound remarkable, and it can also change the voice, invent consonants, or turn a rough take into a slightly different person; it needs an Adobe account and an upload, with a limit on the free tier. This page does the opposite: it keeps your real voice and cleans and polishes it with tools whose effect you can predict, on your own computer. A recording made in a bathroom will still sound like a bathroom, a little less so; a recording made in a reasonable room with a fan running and an uneven speaker will sound properly finished. If you need the studio regeneration, Adobe's tool is the one for that.
Getting the best result
Choose -16 LUFS for a podcast (Apple and Spotify both play speech at about that level) and -14 for YouTube. Leave AI noise removal on unless the recording is already clean, in which case untick it: on a studio voice it has nothing to remove and can only take a little air away. Enhance before you edit if the levels are all over the place, or after editing if you are assembling several takes, so they all end up at the same loudness. For a video, the picture is copied untouched and the sound stays in sync, so you can use the result straight away. To cut out long pauses as well, run the result through auto-cut silence; to transcribe it, audio to text reads a cleaned recording more accurately.
What it can't do
It is for speech. On music, the denoiser treats instruments as noise and the EQ is wrong for them; for a song, use the EQ and the loudness normalizer separately. Other people talking in the background are voices, so the denoiser keeps them. It does not remove echo or reverb from a big room, does not fix distortion from a microphone that was too close or too loud, and cannot bring back words that the noise drowned out completely. And it does not change a voice into a studio voice: for that kind of regeneration, see the comparison above.
FAQ
Is this a free alternative to Adobe Podcast Enhance?
For cleaning up a real voice, yes: AI noise removal, EQ, compression and loudness in one click, free, with no account and no upload. Adobe's tool regenerates the voice with a generative model, which this does not; it keeps your voice and polishes it.
Is my recording uploaded?
No. Everything runs in this browser tab. The only download is the AI denoiser itself, about 5 MB, once.
Will it make my voice sound like a different person?
No. It does not re-synthesise the voice; it removes noise and adjusts tone and level, so it is still your voice, cleaner and more even.
Can I enhance the sound of a video?
Yes. Drop the MP4, MOV or WebM and, with "keep the picture" ticked, you get the video back with the enhanced sound, the picture copied untouched and in sync.
What loudness should I choose?
-16 LUFS for podcasts and spoken audio, -14 LUFS for YouTube. The true peak is held at -1 dBTP either way, so it does not clip when the platform converts it.
Does it work on music?
No, it is built for speech: the AI denoiser treats instruments as noise. For music, use the EQ and the loudness normalizer instead.