The Rundown AI homepage

Independent tool overview

Seed Audio 1.0 at a glance

Seed Audio 1.0 is ByteDance's active full-scene audio-generation model for producing speech, dialogue, sound effects, ambience, and other scene audio from one prompt. The BytePlus API costs $0.15 per generated minute, includes 60 trial minutes, supports outputs up to 120 seconds, and accepts text alone, up to three audio references, or one image reference.

Visit the official Seed Audio 1.0 site ↗
Seed Audio 1.0 product preview
Status
Active through the BytePlus API
Best for
Dialogue-led, full-scene audio generation
Pay-as-you-go
$0.15 per generated minute
Free trial
60 generated minutes after activation
Maximum output
120 seconds per request
Languages
20+ language and locale options

Overview

What Seed Audio 1.0 is

Seed Audio 1.0 is designed to generate a coherent audio scene instead of treating narration, character voices, sound effects, and ambience as separate jobs. A prompt can describe the speakers, dialogue, delivery, setting, and surrounding sounds, while millisecond timestamps provide direct control over when dialogue lines begin.

The current BytePlus API exposes text-only, reference-audio, and reference-image generation. Creators can guide the result with an authorized voice or audio clip, a supported TTS or cloned speaker, or one image. The model can return word-level subtitle timing and WAV, MP3, PCM, or Ogg Opus audio, which makes it usable in automated production pipelines as well as one-off creative work.

This is broader than conventional text-to-speech, but it should not be mistaken for ByteDance's dedicated Seed-Music model. Seed Audio 1.0 can compose speech, effects, ambience, and other scene elements together; however, ByteDance lists finer timing control for sound effects, ambience, and music as future work. Teams that primarily need songs or detailed musical structure should compare a specialist music generator.

Use cases

Who Seed Audio 1.0 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Narrative video audio

Create dialogue, vocal delivery, ambience, and scene-level effects for short videos from a unified brief.

Games and interactive stories

Generate character exchanges and atmospheric audio with controllable dialogue timing and consistent voice direction.

Audiobooks and podcasts

Produce expressive multi-character narration, then use returned subtitle timing to align text or downstream edits.

Localization prototypes

Test dialogue-led audio in more than 20 supported language and locale options before a human linguistic and cultural review.

Audio API workflows

Automate two-minute-or-shorter assets with reference inputs, multiple output formats, subtitle data, and watermark controls.

Capabilities

Core Seed Audio 1.0 features

1

Full-scene generation

Combines speech, multiple speakers, sound effects, ambience, and other scene audio in one generated result.

2

100-millisecond dialogue timing

The launch describes prompt-level dialogue timing control in 100 ms intervals for coordinating scripted lines.

3

Text, audio, or image guidance

Generate from text alone, guide with up to three audio clips, select a supported or cloned speaker, or use one reference image. Audio and image references cannot be mixed in the same request.

4

Voice direction and consistency

Describe a voice in text, use an authorized reference, or combine voice direction with references to guide timbre, emotion, rhythm, and speaking style.

5

20+ languages and locales

The API lists English, Chinese, Japanese, Korean, European and Latin American languages, and several Southeast Asian languages and locales.

6

Two-minute generation

The API supports an original model output of up to 120 seconds, while ByteDance also describes continuation for longer creative sequences.

7

Subtitles and format controls

Optionally return word-level timestamps and choose WAV, MP3, PCM, or Ogg Opus with configurable sample rate, speed, pitch, and loudness.

8

Watermark controls

The API can add an audible rhythm marker and implicit metadata to identify AI-generated audio.

Process

How the Seed Audio 1.0 workflow works

  1. Step 1

    Write the scene as a script

    Specify speakers, exact dialogue, vocal direction, environment, effects, pacing, and the intended length instead of relying on a vague mood prompt.

  2. Step 2

    Choose one reference mode

    Use text only, up to three audio references, a speaker ID, or one image. Obtain permission for every voice, clip, image, and character likeness before uploading it.

  3. Step 3

    Place dialogue timing

    Add millisecond start times when lines must align with a visual edit, then leave room for pauses, effects, and ambience.

  4. Step 4

    Generate a short test

    Validate pronunciation, voice identity, timing, scene balance, and prompt adherence with a short sample before generating the full asset.

  5. Step 5

    Review and finish externally

    Listen for artifacts, unwanted impersonation, incorrect language, clipped speech, timing drift, and unsafe content, then mix or edit the approved output in an audio editor.

Cost

Seed Audio 1.0 pricing and free plan

BytePlus bills Seed Audio 1.0 by the model's original generated duration with second-level precision. Pay-as-you-go is $0.15 per minute, and service activation includes 60 trial minutes. Prepaid packages reduce the effective minute price but expire after one year.

Free trial

60 generated minutes

A one-time usage allowance provided when the Audio 1.0 service is activated.

  • Available after service activation
  • Measured by generated audio duration
  • Use the allowance to test representative languages, voices, and scene types
  • BytePlus does not describe this as a recurring monthly allowance

Pay-as-you-go

$0.15 per minute

Usage-based API billing with no prepaid volume commitment.

  • Equivalent to $0.0025 per generated second
  • Billed using the model's original output duration
  • Billing precision is one second
  • A maximum-length 120-second result costs about $0.30

Prepaid — 200 minutes

$28.50

Small prepaid package with a one-year validity period.

  • Effective rate: $0.1425 per minute if fully used
  • Expires after one year
  • Savings depend on consuming the full allowance

Prepaid — 200,000 minutes

$27,000

High-volume annual package for production workloads.

  • Effective rate: $0.135 per minute if fully used
  • Expires after one year
  • Requires a substantial upfront commitment

Prepaid — 500,000 minutes

$64,000

Larger annual package with a lower effective minute rate.

  • Effective rate: $0.128 per minute if fully used
  • Expires after one year
  • Model acceptance rate and re-generation volume should be included in capacity planning

Prepaid — 2,000,000 minutes

$240,000

The largest package listed on the public pricing page.

  • Effective rate: $0.12 per minute if fully used
  • Expires after one year
  • Best evaluated against an actual annual usage forecast

Pricing checked . Check current pricing at the source ↗

Assessment

Seed Audio 1.0 strengths and limitations

Where it stands out

  • One model can compose dialogue, sound effects, ambience, and other scene audio together.
  • Dialogue timing can be specified at 100 ms intervals.
  • Text descriptions and authorized reference audio provide flexible voice direction.
  • The API supports more than 20 languages and locale variants.
  • Text-only, audio-reference, speaker, and image-reference workflows are available.
  • Word-level subtitle timestamps simplify synchronization and review.
  • Multiple output formats, sample rates, speed, pitch, and loudness controls support production integration.
  • Pricing, a free trial allowance, package sizes, and billing precision are publicly documented.

What to consider

  • A single API request is capped at 120 seconds of original model output.
  • Fine-grained timing control currently focuses on character dialogue; ByteDance lists detailed timing for effects, ambience, and music as future work.
  • The model is not a substitute for Seed-Music or another specialist system when song structure and musical control are the primary job.
  • Audio and image references cannot be combined in the same request.
  • Reference audio is limited to three clips, 30 seconds and 10 MB each; image guidance is limited to one file up to 10 MB.
  • ByteDance's usability and mean-opinion-score results are vendor-reported and do not predict performance on a specific script, accent, or production environment.
  • Multilingual output still needs native-speaker review for pronunciation, meaning, register, and cultural fit.
  • Voice references create consent, impersonation, copyright, and disclosure risks; only authorized material should be used.
  • The returned audio URL expires after two hours, so production systems must save approved assets promptly.
  • Prepaid packages expire after one year, which can erase the nominal discount if the allowance is not fully used.

Compare

Seed Audio 1.0 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Content Creator

ElevenLabs Image & Video

Choose ElevenLabs when a broader creator platform combining voice, music, sound effects, images, and video is more useful than a focused API.

Explore ElevenLabs Image & Video

Content Creator

Stable Audio 3.0

Consider Stable Audio 3.0 for an open-weight, fully licensed audio-model family and music-oriented generation workflows.

Explore Stable Audio 3.0

Content Creator

Suno AI

Use Suno when complete songs, vocals, lyrics, and music creation are the core requirement.

Explore Suno AI

Questions

Seed Audio 1.0 FAQs

What is Seed Audio 1.0?

Seed Audio 1.0 is ByteDance's full-scene audio-generation model for creating dialogue, speech, sound effects, ambience, and other scene audio from one prompt. It is available through BytePlus.

How much does Seed Audio 1.0 cost?

The BytePlus API costs $0.15 per generated minute, billed to the second using the model's original output duration. Activating the service includes 60 trial minutes, and one-year prepaid packages start at $28.50 for 200 minutes.

How long can Seed Audio 1.0 generate?

The current API supports up to 120 seconds of original model output in one request. ByteDance also describes continuation as part of the model's longer-form workflow.

Does Seed Audio 1.0 generate music?

It can create full audio scenes that include speech, effects, ambience, and other elements, but ByteDance lists finer music timing control as future work and offers Seed-Music as a separate specialist model. Use a dedicated music generator when songs or detailed composition are the main objective.

Can Seed Audio 1.0 clone a voice?

The API can use a cloned speaker or authorized reference audio to guide a voice. Only use voices and recordings you have permission to use, and review outputs for impersonation and disclosure risks.

What languages does Seed Audio 1.0 support?

The current BytePlus API lists more than 20 language and locale options, including English, Chinese, Japanese, Korean, French, German, Italian, Russian, several Spanish and Portuguese locales, and multiple Southeast Asian languages.

Can Seed Audio 1.0 use reference files?

Yes. A request can use up to three audio references, a supported or cloned speaker, or one image reference. Image and audio references cannot be mixed in the same request.

Does Seed Audio 1.0 return subtitles?

Yes. The API can return utterance- and word-level subtitle timing when subtitle output is enabled.

Bottom line

Our Seed Audio 1.0 verdict

Seed Audio 1.0 is a compelling API for short, dialogue-led scenes that would otherwise require separate voice, effect, and ambience passes. Its $0.15-per-minute list price is transparent, and the 60-minute trial is enough for a meaningful acceptance-rate test. The important boundary is control: dialogue timing is the mature focus today, while detailed control over effects, ambience, and music remains on ByteDance's roadmap. Use it for scene composition, not as an automatic replacement for a dedicated music model or final human audio review.

Visit Seed Audio 1.0 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.