Promptable source separation

Pull one sound out of the mix just by describing it

Type what you want to hear, click it in the picture, or mark when it happens. SAM Audio Tool separates that sound from the rest of the recording and hands you both halves.

Works on plain audio files and on the audio track of a video.

3
ways to prompt: text, click, timespan
2
stems back every run: target and residual
1
workspace for sound, music and speech
street-interview.mp4Separating
Prompt

the acoustic guitar, not the crowd

Target — guitar
Residual — everything else

Illustration of the separation view. Both stems are exportable.

What it does

One prompt, one sound, cleanly lifted out

Separation used to mean picking from a fixed menu of stems. Here you say what you are after, in the way that suits the recording.

Text prompts

Describe the target in plain language — "the saxophone", "the woman answering the question", "the rain on the window" — and it is lifted out of the mixture.

Visual prompts

When the source is on screen, click it. Pointing at the instrument, the speaker or the passing car is often faster and less ambiguous than writing a description.

Span prompts

Drag across the seconds where the sound is clearly audible. The selection becomes the reference, which is ideal for a noise you cannot name.

Mixed prompting

Combine the three. A click for who, a phrase for what, and a timespan for when — all in a single request, so hard cases stop being guesswork.

Target and residual

Every run returns two files: the sound you asked for, and the full recording with that sound taken out. Keep either one, or work with both.

Audio and video in

Drop in a WAV, an MP3 or a video file. When there is a picture to look at, it is used to make sense of what is making the sound.

Across the whole spectrum

Sound, music and speech in one place

Most tools specialise in one of the three. Pick a domain to see the kind of work each one covers.

Everyday sound, lifted out of the noise

Field recordings, room tone, machinery, wildlife, traffic, doors, footsteps — the long tail of sound that never fits a preset category. Describe it and pull it out.

  • Isolate a single event from a busy ambience
  • Strip an intrusive noise while keeping the take
  • Build clean references from location recordings
Isolated targetGeneral sound
Residual mix

Waveforms shown are illustrative. Your own result depends on the recording and the prompt you write.

The workflow

Four steps from a messy recording to two clean stems

No session setup, no routing, no plugin chain. The prompt is the whole interface.

  1. 01

    Load the recording

    Upload an audio file, or a video when the source is visible on screen. The timeline and waveform are built for you.

  2. 02

    Say what to isolate

    Write a short description, click the source in the frame, or drag over the seconds where it is audible. Combine them when the mix is difficult.

  3. 03

    Listen and adjust

    Audition the target against the residual. If the wrong thing came out, rewrite the prompt or point somewhere more specific and run it again.

  4. 04

    Export both halves

    Download the isolated sound and the remainder as separate files, ready for your editor, your DAW or the next step in the pipeline.

Prompt gallery

What a good prompt looks like

Prompts read like an instruction to an engineer, not like a search query. Say what to keep, or what to lose — both work.

Text

the seagull, not the wind

Target: the bird call. Residual: the coastal ambience it was sitting in.

Music

lead vocal only, no harmonies

Target: the front vocal line. Residual: the full backing, harmonies included.

Visual

the person in the red jacket speaking

Target: that speaker's dialogue. Residual: the other voices and the street.

Span

whatever is making this sound at 0:12–0:19

Target: the source heard in that window. Residual: the rest of the take.

Removal

take out the passing traffic

Residual becomes the deliverable; the traffic is exported separately.

Mixed

the phone ringing behind the dialogue

Target: the ringtone. Residual: clean dialogue you can cut with.

Short and specific beats long and vague. If two sources are similar, add where or when.

Where it earns its keep

Built for anyone who works with recorded sound

Wherever a recording contains more than you wanted, a prompt is faster than a rescue edit.

Podcasts and interviews

Rescue a good answer recorded in a bad room, or separate two guests who talked over each other on one microphone.

Film and video post

Recover location dialogue, lift a sound effect out of a scratch track, or clear a band from a shot you cannot reshoot.

Music practice and study

Solo the part you are learning, mute it to play along, or take a phrase apart to work out what is actually going on.

Accessibility and hearing

Bring a single voice forward and push competing sound back, so a recording stays followable for the person listening.

Research and datasets

Build labelled examples from real mixtures instead of synthetic ones, and keep the residual as a matched counterpart.

Short-form and social

Keep the moment, drop the mess. Pull clean audio out of a phone clip before it goes into the edit.

Questions

Frequently asked

The short answers to what people ask before their first separation.

Traditional separation tools split a track into a fixed set of stems — vocals, drums, bass, other. Promptable separation lets you define the target yourself at request time: you describe the sound, click it on screen, or mark when it happens, and that becomes the thing that gets isolated. The vocabulary is not limited to a preset list.

Describe the sound. Keep the one you want.

Start with a recording you have already given up on and see what a single sentence pulls out of it.

Start separating

No install, no session setup — bring a file and a prompt.