The Rundown AI homepage

Independent tool overview

SAM Audio at a glance

SAM Audio is Meta's prompted audio-separation model for isolating a target sound and everything left behind. It supports natural-language, visual, and time-span prompts across speech, music, and general sound, with a public playground and downloadable checkpoints for technical users.

Visit the official SAM Audio site ↗
SAM Audio product preview
Developer
Meta AI
Primary task
Prompted audio separation
Prompt types
Text, visual, and time span
Access
Playground and gated checkpoints

Overview

What SAM Audio is

SAM Audio applies the Segment Anything idea to sound. Give it a mixed audio or audiovisual source and describe a target such as “dog barking,” click the visible source in a video, or mark a time span where the sound appears. The model generates a target stem and a residual stem containing the remainder.

Unlike tools built only for vocals, instruments, or speech cleanup, SAM Audio is intended as a general foundation model. Meta released small, base, and large checkpoints plus variants tuned for target correctness and visual prompting, along with code, research evaluations, and an interactive Segment Anything Playground.

The downloadable version is aimed at developers and researchers. Checkpoint access is gated through Hugging Face, Python 3.11 or later is required, and Meta recommends a CUDA-compatible GPU. It is source-available under Meta's custom SAM License, not a conventional hosted editing subscription.

Use cases

Who SAM Audio is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

General sound isolation

Separating flexible real-world targets such as traffic, animals, effects, machinery, or ambience without relying on a fixed stem list.

Audio cleanup research

Testing background-noise and unwanted-event removal with text or time-span prompts.

Music and speech experiments

Exploring instrument, vocal, speaker, and speech separation with one unified model family.

Audiovisual workflows

Using a click or visual mask on a sounding object or person in video to guide the extracted audio.

Capabilities

Core SAM Audio features

1

Text prompting

Describe the desired source with a concise noun or verb phrase, such as “man speaking” or “car honking.”

2

Visual prompting

Provide video frames and a mask around a visible sound source to isolate the audio associated with it.

3

Span prompting

Mark positive time ranges where the target occurs, giving the model a temporal example of the sound to separate.

4

Target and residual stems

Each separation returns the isolated target and a complementary track containing everything else.

5

Automatic span prediction

For text prompts, the model can predict likely target spans before separation, which is useful for intermittent sound events.

6

Candidate reranking

Generate multiple separation candidates and choose among them with CLAP, SAM Audio Judge, or ImageBind scoring, trading more compute for potentially better output.

7

Six checkpoint variants

Meta provides small, base, and large models, plus corresponding TV variants focused on target correctness and visual prompting.

Process

How the SAM Audio workflow works

  1. Step 1

    Choose an input and target

    Start with audio or video you have the right to process, then identify the exact sound, speaker, instrument, or event to isolate.

  2. Step 2

    Select a prompt mode

    Use a short text description, a visual mask, a time span, or a combination that supplies the clearest evidence.

  3. Step 3

    Run separation

    Try the public playground or load an approved checkpoint into Meta's Python inference code on compatible hardware.

  4. Step 4

    Compare target and residual

    Listen for missing target detail, bleed, phase changes, and artifacts in both output stems.

  5. Step 5

    Tune for the use case

    Adjust span prediction and candidate reranking only when the quality gain justifies additional latency and GPU memory.

Cost

SAM Audio pricing and free plan

Meta does not publish a paid SAM Audio plan or managed API rate. The playground and downloadable research materials have no listed usage price; self-hosting requires your own compute and acceptance of the custom SAM License.

Segment Anything Playground

No price listed

A browser demo for trying SAM Audio on provided or uploaded media.

  • Interactive prompt testing
  • No production SLA
  • Usage availability may change

Self-hosted checkpoints

No Meta license fee listed

Download gated model checkpoints and run inference with Meta's research code under the SAM License.

  • Hugging Face access approval required
  • Python 3.11 or later
  • CUDA-compatible GPU recommended
  • Hosting and engineering costs are separate

Managed API

Not offered

Meta does not list a production SAM Audio API or per-minute hosted price.

  • No official API rate card
  • No managed service tier
  • Teams must build or source their own deployment

Pricing checked . Check current pricing at the source ↗

Assessment

SAM Audio strengths and limitations

Where it stands out

  • One model family covers general sound, speech, speakers, music, and instruments
  • Text, visual, and temporal prompts give users multiple ways to identify a target
  • Returns both the isolated source and the residual mixture
  • Public code, checkpoints, examples, evaluation tools, and research paper support deeper testing
  • Candidate generation and reranking expose an explicit quality-versus-compute control

What to consider

  • SAM Audio is a research model and codebase, not a turnkey editor with a production support agreement.
  • Checkpoint access is gated, local setup is technical, and Meta recommends CUDA hardware.
  • Generating and reranking more candidates can improve results but increases latency and memory use.
  • Separation quality depends on the source, prompt, overlap, and model variant; bleed and audible artifacts still require human review.
  • The model uses Meta's custom SAM License, so organizations should review its use and redistribution terms before deployment.

Compare

SAM Audio alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

Auphonic

Auphonic is a managed production tool for podcast and spoken-word cleanup, leveling, noise reduction, and delivery rather than open-ended source separation.

Explore Auphonic

Content Creator

Descript

Descript combines transcript-based audio and video editing with creator-friendly cleanup tools for people who do not want to self-host a research model.

Explore Descript

Marketing

ElevenLabs

ElevenLabs offers hosted voice and audio tools for production workflows, including a broader commercial platform and API ecosystem.

Explore ElevenLabs

Questions

SAM Audio FAQs

What does SAM Audio do?

It isolates a requested sound from mixed audio or video and returns two outputs: the target sound and the residual containing everything else.

How can I prompt SAM Audio?

Use a short text description, a visual mask on a sound-producing object in video, a marked time span where the sound occurs, or a combination of those signals.

Is SAM Audio free?

Meta publishes the research code and gated checkpoints without a listed license fee and provides a public playground. You still pay for your own compute, engineering, storage, and any production infrastructure, and use is governed by the SAM License.

Does SAM Audio have an API?

Meta does not currently list an official managed SAM Audio API or per-minute price. Developers can run the released checkpoints through the project code after receiving access.

What hardware does SAM Audio need?

The official repository requires Python 3.11 or later and recommends a CUDA-compatible GPU. Actual memory and speed depend on model size, input, span prediction, and the number of reranking candidates.

Bottom line

Our SAM Audio verdict

SAM Audio is unusually flexible for researchers and product teams that need more than fixed vocal or instrument stems. The playground is the easiest way to test prompt quality, while self-hosting makes sense only for teams comfortable with gated checkpoints, GPU inference, a research-oriented codebase, and Meta's custom license. Production buyers seeking reliable cleanup with support should compare managed audio tools first.

Visit SAM Audio website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.