Bottom Line Up Front (BLUF): The Vocal Separator splits a track into two stems — vocals and instrumental — right in your browser. Upload a song, press the button and get the vocals on their own and a karaoke backing track: both stems can be played on the page and downloaded as WAV files. The file is never uploaded, everything runs on your own device, no tokens are spent and there is no paid API. A video file works too — the tool takes the audio track from it.
This is music source separation. You feed in a song and get two stems back: the vocals on their own and the instrumental — a backing track without the voice. It helps when you need a karaoke backing track, when you want to cut the vocals out and sing over it, build a remix, or simply hear how the arrangement sounds without the singer.
The isolated voice is useful on its own too: as a reference, in an edit, or to study the phrasing.
The model breaks the track into four parts — drums, bass, other and vocals. You are shown two stems: the vocals are taken as they are, and the instrumental is built from all the other parts without the voice. That way the backing track sounds whole instead of like fragments of separate instruments.
The result is delivered as WAV files at 44.1 kHz stereo — play them right on the page and download them.
Your track never leaves your computer: neither the file nor the finished stems are uploaded anywhere. The model is loaded into the browser and runs on your hardware — the computation goes through WebGPU, and if the graphics card is unavailable the engine honestly falls back to the CPU and warns you: it will be noticeably slower.
No tokens are spent on separation: this is a free tool that runs on your device, not on our server.
Audio: MP3, WAV, M4A, OGG, FLAC and other common formats. You can also upload a video file — the tool takes its audio track.
The tool works best in a recent Chrome, Edge or Safari. In a browser without WebGPU the separation falls back to the CPU: it still works, just slower.
On the first run the browser downloads the model weights once — about 172 MB. They stay in the browser cache, and later runs start immediately.
Time depends on the length of the track and the power of your device: a short fragment separates quickly, a long song takes longer. The progress bar shows where it is, and you can cancel the run.
Correct. The model and its weights are loaded into the browser and the computation runs on your device. The track and the finished stems are never uploaded — the tool has no paid API at all.
Nothing. No tokens are spent: this is a free tool that runs on your computer.
Two: the vocals on their own and the instrumental — a backing track without the voice. Both can be played on the page and downloaded as WAV files.
Yes. The tool takes the audio track from the video and separates it just like a regular track.
The browser downloads the model weights once — about 172 MB. After that they come from the cache and the run starts immediately.
The tool keeps working on the CPU — slower, but with the same result. It warns you that the fallback mode is on.