Lyrics from a song: the words, synced.

Drop a song. The page takes the vocal out of the mix, listens to it, and writes the words line by line with their times: plain text, an LRC file for karaoke and music players, or SRT and VTT captions for a lyric video. All in your browser; the song is never uploaded.

drop a song

MP3, WAV, M4A, FLAC, OGG and most other audio, or an MP4 or MOV. Up to 10 minutes.

At a glance

TakesA song up to 10 minutes: MP3, WAV, M4A, FLAC, OGG, AIFF, or the sound of an MP4 or MOV
HowMDX-Net separates the vocal (the stem splitter's model, 64 MB), then Whisper base transcribes it with the time of each word (135 MB); both downloaded once and kept
GivesLRC synced lyrics, SRT and VTT captions, plain text; every line editable before you save
LanguagesWhisper's 99, detected automatically or chosen
Uploads your songNo

Why the vocal comes out first

Speech recognisers are trained on people talking, not singing over drums, guitars and a bass. Given a full mix they miss words, and in the instrumental breaks they invent them. So the page first runs the same separation model as the stem splitter to lift the vocal off the track, then gives Whisper only the voice. A line is kept only where the separated vocal is actually sounding, which is what stops the recogniser writing lines over an instrumental break.

Separation is the slow part. With a GPU (current Chrome or Edge on most laptops) it runs at about the song's own length; on the CPU it takes roughly ten times that, so without a GPU the box starts unticked. A clear vocal over a light backing, an acoustic song or a podcast jingle, often reads well enough without it.

LRC, SRT or plain text

LRC is the synced-lyrics format: one line per lyric line with its start time, like [01:12.40]And the night goes on. Music players such as foobar2000, MusicBee, Poweramp and Musixmatch-style apps show it in time with the song when the .lrc file sits next to the audio file with the same name, and karaoke apps and lyric-video makers import it. A blank timed line is added where a long instrumental starts, so the last words do not hang on screen. SRT and VTT are caption files: put them on a lyric video with add subtitles to video, or upload them to YouTube with the video. Text is just the words, a line per row, for a lyric sheet or a search.

Correcting the words

Every line is a text box. Click the time beside a line to play the song from there, fix what was misheard, and the downloads use your corrected text with the original timing. Expect to fix some: sung vowels are stretched, words are slurred on purpose, backing vocals overlap the lead, and a recogniser that has never heard the song guesses at made-up words and names. Lines are cut where the singer pauses or a sentence ends; join or split them in a text editor if you want them to match the verses exactly.

What it can't do

It does not look lyrics up: it only knows what it hears, so it cannot tell you the official words, and mumbled or heavily processed vocals (screaming, auto-tune used as an effect, a choir) come out partly wrong. Times are per line, not per syllable, so karaoke apps that colour each word as it is sung will show whole lines. Lyrics belong to their writers; use what you make here for your own listening, learning and practice, or for songs you have the rights to.

FAQ

How do I get the lyrics of a song from the audio?

Drop the song above and press Get the lyrics. The vocal is separated from the music, transcribed by Whisper in your browser, and shown line by line with times. Correct anything misheard, then download LRC, SRT, VTT or text.

What is an LRC file and how do I use it?

An LRC file is lyrics with a time on each line. Save it next to the song with the same name (song.mp3 and song.lrc) and players like foobar2000, MusicBee and Poweramp show the words in time; karaoke and lyric-video apps import it too.

Is my song uploaded?

No. Both models are downloaded to your browser and the song is separated and transcribed there. The only downloads are the models, which are the same for everyone.

How accurate is it?

It depends on how clear the singing is. Clean pop and acoustic vocals come out mostly right; rap, screaming, heavy effects and choirs need more correction. Every line is editable before you download.

Can I make a lyric video with it?

Yes. Download the SRT, then put it on your video with add subtitles to video, which can burn the words into the picture. For a song with no video, make one with MP3 to MP4 first.

Why is it slow without a GPU?

Separating the vocal is a large model. On a GPU it runs at about the song's length, on the CPU about ten times longer. Untick "Separate the vocal first" for a clear vocal to skip it.