Remove background noise from audio and video.
Removes steady background noise from a recording — fan hum, air conditioning, tape hiss, preamp noise — using spectral gating, entirely in your browser with nothing uploaded. It analyses the noise floor automatically, so there is nothing to configure beyond how hard to push it.
drop a noisy recording or video
MP3, WAV, M4A, FLAC, or an MP4, MOV or WebM video. Voice recordings work best.
At a glance
| Method | Short-time Fourier transform with per-band spectral gating |
|---|---|
| Removes well | Fan hum, air conditioning, tape hiss, preamp hiss, mains hum |
| Does not remove | Spectral: anything irregular (coughs, traffic, keyboards). AI: other people's voices, and it is for speech, not music |
| Noise profile | Estimated automatically from the whole file |
| Output format | MP3 (128–320 kbps) or WAV |
| Practical size limit | ~500 MB per file (browser memory) |
| Uploads your file | No — runs entirely in the page |
| Signup required | No |
How it actually works
The recording is chopped into short overlapping frames and each one is converted into a frequency spectrum — a measurement of how much energy sits in each frequency band at that instant. Steady noise has a distinctive signature in that view: it is present in every frame at roughly the same level, because a fan does not stop humming between your words.
The tool estimates that floor by taking a low percentile of each frequency band across the whole file, then goes back through and attenuates any band that sits close to it, leaving louder content untouched. Speech and music, which are far above the floor when present, come through; the constant hum underneath is pulled down by 9 to 24 dB depending on the strength you choose.
The attenuation uses a soft knee rather than a hard cut. A hard gate produces the warbling, underwater artefact that people associate with badly denoised audio — sounds flickering in and out as they cross the threshold. Easing between attenuated and untouched avoids most of that.
What it fixes, and what it can't
It fixes steady noise. That means anything continuous and unchanging: computer fans, air conditioning, refrigerator hum, the 50 or 60 Hz buzz from bad grounding, tape hiss, the hiss of a cheap preamp or a phone microphone turned up too far. These are the things that make an otherwise fine recording sound amateur, and they come out cleanly.
Spectral gating cannot fix irregular noise. A cough, a door closing, a car passing, a keyboard, a dog barking: none of these are stationary, so there is no consistent floor to estimate and nothing for a spectral gate to grip. That is what the AI method is for.
The AI method: any noise behind a voice
Choose "Any noise behind a voice (AI)" and the page runs RNNoise, a small neural network from Xiph (the Opus codec people) trained on thousands of hours of speech and noise to keep a voice and drop everything else. It does not need the noise to be steady: traffic, keyboard clatter, a fan that speeds up, a café's clatter and a dog barking between sentences are all fair game. It runs in your browser like the rest of the page, at about a third of the audio's length on a laptop, and the result lines up with the original to the sample, so a video stays in sync.
Measured on speech under noise bursts and clicks that switch on and off, the kind spectral gating cannot touch: at a realistic level (the voice a little louder than the noise), the pauses came out 43.5 dB quieter and the speech matched the clean original at a correlation of 0.94, against 10 dB and 0.90 for spectral gating. With noise three times louder than the voice, the pauses still came out 48 dB quieter, and the speech went from barely recognisable (0.28) to clearly there (0.69). It is built for voices: it will not clean up music, and it keeps human voices, so other people talking in the background stay. Strength blends the cleaned voice with the original: gentle keeps some room sound, strong is the cleanest.
Choosing a strength
Medium is the right default for speech and handles most room tone and fan noise without touching the voice.
Gentle is the setting for music. Music occupies far more of the spectrum than speech does, including quiet passages that a more aggressive gate would mistake for noise — reverb tails, cymbal decay and the quiet end of a fade are all vulnerable. Gentle removes less noise but is much less likely to dull the recording.
Strong is for recordings that are genuinely noisy: a phone recording in a busy room, an old tape transfer, a video call capture. Expect some cost to the voice — a slightly hollow or processed quality — because at that level of reduction you are removing content that overlaps with the noise.
A useful order of operations: remove the silence first if the recording has long gaps, denoise second, and normalize last. Denoising lowers the overall level, so normalizing afterwards puts it back where you want it.
What you lose
Some high-frequency detail, always. Hiss and the top end of a voice occupy the same region of the spectrum, so pulling one down inevitably takes a little of the other. On speech this reads as very slightly duller and is almost never a problem; on music it is more noticeable, which is why Gentle exists.
Processing time. This is the slowest tool on the site by some margin — the file is transformed to the frequency domain, modified, and transformed back, twice over, and a long recording means a lot of frames. A three-minute file takes a few seconds; a one-hour interview takes a while. The progress bar is real.
You do not lose the original. Denoising is destructive in the sense that the output is a new file, so keep the source until you have listened to the result — particularly on music, where it is worth comparing the two before committing.
FAQ
Can I remove background noise from a video?
Yes. Drop the MP4, MOV or WebM and, with "keep the picture" ticked, you get the video back with the cleaned sound and the picture copied untouched, so nothing is re-compressed and the sound stays in sync. On a test clip the picture came back with every frame and the cleaned sound lined up to the sample (0.0 ms of offset).
What kind of noise can this actually remove?
Steady, continuous noise: fan and air-conditioning hum, refrigerator noise, mains buzz, tape hiss, and the hiss of a cheap microphone or preamp. These have a constant signature that spectral gating can identify and pull down.
Can it remove a dog barking, traffic or keyboard noise?
Yes, with the AI method: choose "Any noise behind a voice (AI)". It keeps the voice and drops irregular noise such as traffic, typing and barking, which spectral gating cannot. Other people talking in the background are voices too, so it keeps them; for those, cutting around them is still the practical answer.
Why does my voice sound slightly hollow afterwards?
You are probably on Strong. At high reduction levels, content that overlaps with the noise in frequency gets removed along with it. Drop to Medium, or Gentle for music.
Why is this slower than the other tools?
Because it transforms the audio into the frequency domain and back rather than just re-encoding it. Every frame is analysed twice. A few minutes of audio takes seconds; a long interview takes noticeably longer.
Should I denoise before or after normalizing?
Denoise first, normalize afterwards. Noise reduction lowers the overall level, so normalizing last puts the result where you want it. If the recording has long gaps, remove the silence before either step.