🤖

Source Separation

Separate audio into individual stems (vocals, drums, bass, other) using Demucs on GPU.

GET 1 credit /v1/audio/separate
curl "https://audio.toolkitapi.io/v1/audio/separate?audio_url=https://example.com/song.mp3&model=htdemucs&stems=4stem"
import httpx

resp = httpx.get(
    "https://audio.toolkitapi.io/v1/audio/separate?audio_url=https://example.com/song.mp3&model=htdemucs&stems=4stem",
)
print(resp.json())
const resp = await fetch("https://audio.toolkitapi.io/v1/audio/separate?audio_url=https://example.com/song.mp3&model=htdemucs&stems=4stem", {
});
const data = await resp.json();
console.log(data);
# See curl example
Response 200 OK
{
  "status": "ok",
  "vocals_url": "https://storage.example.com/audio/vocals.mp3",
  "drums_url": "https://storage.example.com/audio/drums.mp3",
  "bass_url": "https://storage.example.com/audio/bass.mp3",
  "other_url": "https://storage.example.com/audio/other.mp3",
  "task_id": "abc123"
}

Try It Live

Live Demo

Description

Separate audio into individual stems (vocals, drums, bass, other) using Demucs on GPU.

How to Use

1

1. Provide a publicly accessible audio URL.

2

2. Choose a model: `htdemucs` (default, good quality) or `htdemucs_ft` (fine-tuned, better for music).

3

3. Choose stem separation: `2stem` for vocals + accompaniment, or `4stem` for vocals, drums, bass, and other.

4

4. Each stem is returned as a separate downloadable audio URL.

About This Tool

Separate mixed audio into individual stems using Meta's Demucs deep learning model on GPU. Choose between 2-stem (vocals + accompaniment) and 4-stem (vocals, drums, bass, other) separation.

Perfect for creating karaoke tracks, remixing songs, isolating instruments, and audio analysis.

Why Use This Tool

Frequently Asked Questions

What's the difference between the two models?
`htdemucs` is the standard model — fast and good quality. `htdemucs_ft` is fine-tuned specifically for music and may produce cleaner separations at the cost of slightly longer processing.
Can I separate more than 4 stems?
Currently, 4-stem is the maximum. Additional stem types may be supported in future versions.
How long does separation take?
A typical 3-4 minute song processes in 15-30 seconds on GPU. Longer tracks scale linearly.

Start using Source Separation now

Get your free API key and make your first request in under a minute.