Text prompts
Describe the target in plain language — "the saxophone", "the woman answering the question", "the rain on the window" — and it is lifted out of the mixture.
Type what you want to hear, click it in the picture, or mark when it happens. SAM Audio Tool separates that sound from the rest of the recording and hands you both halves.
Works on plain audio files and on the audio track of a video.
“the acoustic guitar, not the crowd”
Illustration of the separation view. Both stems are exportable.
Separation used to mean picking from a fixed menu of stems. Here you say what you are after, in the way that suits the recording.
Describe the target in plain language — "the saxophone", "the woman answering the question", "the rain on the window" — and it is lifted out of the mixture.
When the source is on screen, click it. Pointing at the instrument, the speaker or the passing car is often faster and less ambiguous than writing a description.
Drag across the seconds where the sound is clearly audible. The selection becomes the reference, which is ideal for a noise you cannot name.
Combine the three. A click for who, a phrase for what, and a timespan for when — all in a single request, so hard cases stop being guesswork.
Every run returns two files: the sound you asked for, and the full recording with that sound taken out. Keep either one, or work with both.
Drop in a WAV, an MP3 or a video file. When there is a picture to look at, it is used to make sense of what is making the sound.
Most tools specialise in one of the three. Pick a domain to see the kind of work each one covers.
Field recordings, room tone, machinery, wildlife, traffic, doors, footsteps — the long tail of sound that never fits a preset category. Describe it and pull it out.
Waveforms shown are illustrative. Your own result depends on the recording and the prompt you write.
No session setup, no routing, no plugin chain. The prompt is the whole interface.
Upload an audio file, or a video when the source is visible on screen. The timeline and waveform are built for you.
Write a short description, click the source in the frame, or drag over the seconds where it is audible. Combine them when the mix is difficult.
Audition the target against the residual. If the wrong thing came out, rewrite the prompt or point somewhere more specific and run it again.
Download the isolated sound and the remainder as separate files, ready for your editor, your DAW or the next step in the pipeline.
Prompts read like an instruction to an engineer, not like a search query. Say what to keep, or what to lose — both work.
“the seagull, not the wind”
Target: the bird call. Residual: the coastal ambience it was sitting in.
“lead vocal only, no harmonies”
Target: the front vocal line. Residual: the full backing, harmonies included.
“the person in the red jacket speaking”
Target: that speaker's dialogue. Residual: the other voices and the street.
“whatever is making this sound at 0:12–0:19”
Target: the source heard in that window. Residual: the rest of the take.
“take out the passing traffic”
Residual becomes the deliverable; the traffic is exported separately.
“the phone ringing behind the dialogue”
Target: the ringtone. Residual: clean dialogue you can cut with.
Short and specific beats long and vague. If two sources are similar, add where or when.
Wherever a recording contains more than you wanted, a prompt is faster than a rescue edit.
Rescue a good answer recorded in a bad room, or separate two guests who talked over each other on one microphone.
Recover location dialogue, lift a sound effect out of a scratch track, or clear a band from a shot you cannot reshoot.
Solo the part you are learning, mute it to play along, or take a phrase apart to work out what is actually going on.
Bring a single voice forward and push competing sound back, so a recording stays followable for the person listening.
Build labelled examples from real mixtures instead of synthetic ones, and keep the residual as a matched counterpart.
Keep the moment, drop the mess. Pull clean audio out of a phone clip before it goes into the edit.
The short answers to what people ask before their first separation.
Traditional separation tools split a track into a fixed set of stems — vocals, drums, bass, other. Promptable separation lets you define the target yourself at request time: you describe the sound, click it on screen, or mark when it happens, and that becomes the thing that gets isolated. The vocabulary is not limited to a preset list.
Start with a recording you have already given up on and see what a single sentence pulls out of it.
Start separatingNo install, no session setup — bring a file and a prompt.