Voice changer: live mic or recordings.
Change your voice live through the microphone, or change how a voice sounds in an audio file or a video you already have, such as a voice memo, a WhatsApp voice note, a voice-over, a podcast clip or a phone video: deeper or higher while still sounding like a person, male to female or female to male, or something else entirely, with robot, monster, chipmunk, old radio and phone-line voice effects. Pitch and formants are separate controls, which is what keeps a lowered voice from sounding like a slowed-down tape. Runs in your browser; nothing is uploaded.
drop a voice recording or a video
MP3, WAV, M4A, voice memos, FLAC, or an MP4, MOV or WebM video. Batch supported.
Live voice changer: change your voice as you speak
Hear yourself through the effect in real time and record the result. Use headphones: through speakers the changed voice feeds back into the microphone.
At a glance
| Method | Pitch-synchronous overlap-add (PSOLA) for pitch, sinc resampling for formants, so the two move independently |
|---|---|
| Pitch range | One octave down to one octave up, in half-semitone steps |
| Accuracy | Within 3 cents of the requested pitch on a test voice, with the length unchanged to the sample |
| Presets | Deeper, higher, more masculine, more feminine, chipmunk, monster, robot, alien, old radio, telephone, cave |
| Keeps | The sample rate, channels and length of the original |
| Video | MP4, MOV, MKV or WebM in; the same video back with the new voice, the picture copied without re-encoding (WebM comes back as MP4) |
| Uploads your file | No — everything runs in the page |
Make a voice deeper without the slowed-tape sound
A voice has two separate qualities that people tend to lump together as "how deep it sounds". Pitch is the note: how fast the vocal folds vibrate, around 100–130 times a second for a typical adult male speaking voice and 180–220 for a typical female one. Formants are the resonances of the throat, mouth and nose that turn that buzz into vowels. Their positions depend on the size and shape of the vocal tract, which is why a child and an adult singing the same note still sound nothing alike.
Speeding up or slowing down a recording moves both by the same amount. That is the chipmunk effect, and it is also why most "deep voice" apps sound like a slowed tape rather than a large person. This tool cuts the voice into slices one vocal-fold cycle long and lays them back down closer together or further apart. That changes the pitch without changing the slices themselves, so the formants stay put. The formant slider then moves the resonances on their own, by resampling before the slices are cut. Lower both and you get a bigger, darker person; lower just the pitch and you get the same person speaking lower; raise just the formants and the voice sounds younger or smaller at the same note.
Male to female, female to male: getting a convincing result
For a believable deeper or higher voice, stay within about four semitones and move the formant a little in the same direction, about a third as far. The Deeper and Higher presets do exactly that. More masculine and More feminine go further on both, which is roughly the difference between typical adult voices. They will not turn one person into a convincing other person, and they are not meant to. Past six or seven semitones every pitch shifter starts to sound processed, because real voices change the way they are produced, not just their note, when they go that far.
Use Preview 8 s to hear the current setting on the start of your file before processing all of it. Clean, dry recordings shift best. Background music or a second voice confuses the pitch detection, so separate the voice first with the AI vocal remover, and reduce hiss with noise reduction before rather than after.
Robot, chipmunk, monster and radio voice effects
Chipmunk raises the pitch nine semitones and the formants by about a third, the sped-up cartoon sound, but without changing the length. Robot flattens the voice to a single note (110 Hz) and adds ring modulation, which is the classic science-fiction approach: the words stay intelligible but the melody of speech disappears. Monster drops the pitch ten semitones and the formants by about a quarter, then adds drive and a dark room. Alien raises both slightly and rings the voice against a 180 Hz tone. Old radio and Telephone cut the voice down to the narrow band those devices carry: roughly 300 Hz to 3.4 kHz for a phone line, with an 8 kHz sample-rate grit on top. Cave adds a long reverb with a tail, so the file gets a few seconds longer.
Every preset sets the sliders when you choose it, and the sliders can then adjust its pitch and formant. A "monster, but less" is one slider move.
Live voice changer: speak and hear it changed
The live section takes your microphone and changes it as you talk, with about 40 milliseconds of delay from the effect plus your device's own audio latency. Wear headphones; through speakers the changed voice goes back into the microphone and howls. Deeper, higher, chipmunk and monster move the pitch; robot and alien ring the voice against a tone; old radio and telephone squeeze it into the narrow band those devices carry. The pitch slider sets any amount from an octave down to an octave up, and the level bar shows the microphone is being heard. Record saves exactly what you hear as an MP3, so it is also a quick way to make a funny voice message or a character line without editing. On a 200 Hz test tone the Deeper setting recorded at 158.7 Hz, which is four semitones down to the tenth of a cent.
Change the voice in a video
Drop a video and, with "keep the picture" ticked, you get the video back with the new voice: the picture stream is copied across untouched, not re-encoded, so it loses nothing and the job takes seconds longer rather than minutes. Because the changed voice is the same length as the original to the sample, it stays in sync with lips and cuts. On a test clip the picture came back with every one of its 125 frames, the length matched to the hundredth of a second, and the Deeper preset measured 394 cents lower against the 400 it asks for. Background music in the video is shifted along with the voice, so for a clip with a soundtrack, separate the voice first with the AI vocal remover. Untick the box to get just the changed audio as MP3, WAV, FLAC or M4A.
What it can't do
This is a voice changer, not a voice cloner: it changes the character of the voice in the recording but cannot make it sound like a specific other person. Desktop tools such as RVC and VoiceStudio do that with neural models trained per voice. The live mode is rougher than the file mode: shifting a voice as it arrives cannot see whole cycles of the waveform, so it cannot hold the formants, and a deeper voice also sounds bigger; for the most natural result, record first and change the file. A browser tab cannot be chosen as a microphone by another app, so the live mode does not plug straight into Discord, Zoom or a game; a desktop app such as Voicemod sits between the microphone and other programs, or the tab's sound can be routed through a virtual audio cable (below). If you need a voice that was never recorded at all, the text to speech page generates one from typed text.
The output keeps the original's sample rate, channel count and length to the sample, except for the cave preset, whose echo adds a tail. Sent from the audio editor, the changed clip comes back lined up exactly where it was.
FAQ
Is it free, and is my recording uploaded?
Free, no signup, and nothing is uploaded. The processing runs in your browser tab, so the recording never leaves your computer.
Can I change the voice in an audio file I already recorded?
Yes, that is what this page is for. Drop in an MP3, WAV, M4A, FLAC, OGG or Opus file (an iPhone voice memo or a WhatsApp voice note works), pick a preset, and download the result as MP3, WAV, FLAC or M4A. Several files can be changed in one go.
How do I make my voice deeper without it sounding slowed down?
Choose the Deeper preset, or lower the pitch by three or four semitones and the formant by about one. The pitch moves without the formants, which is what keeps it sounding like a person rather than a slowed tape.
Can it make a male voice sound female, or the reverse?
The More feminine and More masculine presets move pitch and formants by roughly the difference between typical adult voices. The result is noticeably different, but speaking style matters as much as pitch, so it will not fully convince on its own.
Can I make my voice deeper in a video?
Yes. Drop the MP4, MOV or WebM, choose Deeper (or any preset), and the video comes back with the changed voice and the picture copied untouched, still in sync. MOV and MKV keep their format; WebM comes back as MP4.
Does it work on singing?
Yes, on a solo voice. For a voice in a song, separate it first with the AI vocal remover, change it, and mix it back in the audio editor.
Does it change the length of the file?
No. The output is the same length as the original to the sample, so it still lines up with video. The one exception is the cave preset, which adds its echo tail at the end.
Can I change my voice live, as I talk?
Yes. In the live section, press Start the microphone, pick a voice and speak: you hear the changed voice in your headphones, and Record saves it as an MP3. Measured on a test tone, the Deeper voice lands exactly four semitones down.
Can I use it on Discord, Zoom or in a game?
Not directly: other apps cannot pick a browser tab as their microphone. The general way round that is a virtual audio cable (VB-CABLE on Windows, BlackHole on a Mac): send the browser's sound to the cable and choose the cable as the microphone in the other app. We have not tested every app that way. A desktop voice changer such as Voicemod is the simpler route for calls and games.