One click runs a real source-separation model on your device (about 67 MB, size shown before the one-time download, cached for next time) - no account, no upload, no queue.
Vocal Remover - Make Instrumental & Karaoke Tracks On-Device
Load a song and this page splits it into two downloads - an instrumental (karaoke) track and a vocals-only track - entirely on your device. Nothing is uploaded: the separation model (about 67 MB, fetched once on your click and cached) runs in your browser, so the song never leaves your computer.
Two ways to remove the vocals, and when each one fits
The AI button runs a real source-separation model over the whole song. It works on any mix - mono or stereo, voice panned anywhere - and it returns both the instrumental and the isolated vocals. The tradeoff is time: everything runs locally on one processor thread, so expect roughly 3 to 5 times the song length. A 3-minute track takes about 10 to 15 minutes on a typical laptop, with a live progress line and a time estimate so you always know it is still working.
The Quick voice-cancel button is the classic karaoke trick: it cancels whatever is panned dead-center in a stereo file. It is instant and needs no download at all, but it only helps when the voice actually sits in the center of a stereo mix, and it also thins out anything else that is centered - most often the bass and the kick drum. If you load a mono file, the page tells you plainly that there is no center channel to cancel instead of handing you silence.
| Mode | Download | Speed | Works on | Outputs |
|---|---|---|---|---|
| AI separation | 67 MB, once | 3-5x song length | any file, mono or stereo | instrumental + vocals |
| Quick voice-cancel | none | instant | stereo with centered voice | instrumental only |
What to expect from the result
The example at the top of the page is real output from the same model, prepared ahead of time: the centered voice in a test mix dropped by about 24 decibels while the backing moved less than 1 decibel. On real songs the quality depends on the mix - one small on-device model will not match the big commercial stem services on dense, loud masters, but on typical pop, acoustic, or podcast-style material the vocal drop is clearly audible and the backing stays intact. Both results download as stereo 44.1 kHz WAV files you can drop straight into an editor.
Working with recorded speech instead of music? The audio denoiser removes steady background hiss from a voice recording, and the silence remover cuts the dead air out of it - both run on-device the same way this page does.
Frequently Asked Questions
Is my song uploaded anywhere?
No. The file is decoded and separated entirely inside your browser tab. The only download is the separation model itself (about 67 MB, fetched once from a public model host and cached); your audio never leaves your device on either mode.
Why does the AI separation take so long?
The model processes the song in roughly 6-second passes on a single processor thread, which works out to about 3 to 5 times the song length on a typical laptop. The status line shows which pass it is on and how many minutes are left, so a long run is visible progress, not a hang. Keep the tab open until it finishes.
What is the difference between the AI separation and the Quick voice-cancel?
The AI separation understands what a singing voice sounds like and removes it from any mix, returning both an instrumental and a vocals-only track. The Quick voice-cancel is instant and needs no download, but it simply cancels the center of a stereo mix - it only works when the voice is panned dead-center, it thins centered bass too, and it cannot produce a vocals-only track.
What files can I load, and what do I get back?
Anything your browser can decode: MP3, WAV, M4A, OGG, OPUS, WebM, FLAC, or AAC, up to 10 minutes long. Both results come back as stereo 44.1 kHz 16-bit WAV files, with in-page players so you can compare the original, the instrumental, and the vocals before downloading.
Will the instrumental sound as clean as a paid stem service?
Not always. This is one small on-device model, not a server ensemble - on dense, loud masters some vocal residue can remain. In the prepared example on this page the centered voice dropped by about 24 dB while the backing moved less than 1 dB; typical pop, acoustic, and podcast material separates clearly.