Narrative video audio
Create dialogue, vocal delivery, ambience, and scene-level effects for short videos from a unified brief.
Independent tool overview
Seed Audio 1.0 is ByteDance's active full-scene audio-generation model for producing speech, dialogue, sound effects, ambience, and other scene audio from one prompt. The BytePlus API costs $0.15 per generated minute, includes 60 trial minutes, supports outputs up to 120 seconds, and accepts text alone, up to three audio references, or one image reference.
Visit the official Seed Audio 1.0 site ↗
Overview
Seed Audio 1.0 is designed to generate a coherent audio scene instead of treating narration, character voices, sound effects, and ambience as separate jobs. A prompt can describe the speakers, dialogue, delivery, setting, and surrounding sounds, while millisecond timestamps provide direct control over when dialogue lines begin.
The current BytePlus API exposes text-only, reference-audio, and reference-image generation. Creators can guide the result with an authorized voice or audio clip, a supported TTS or cloned speaker, or one image. The model can return word-level subtitle timing and WAV, MP3, PCM, or Ogg Opus audio, which makes it usable in automated production pipelines as well as one-off creative work.
This is broader than conventional text-to-speech, but it should not be mistaken for ByteDance's dedicated Seed-Music model. Seed Audio 1.0 can compose speech, effects, ambience, and other scene elements together; however, ByteDance lists finer timing control for sound effects, ambience, and music as future work. Teams that primarily need songs or detailed musical structure should compare a specialist music generator.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Create dialogue, vocal delivery, ambience, and scene-level effects for short videos from a unified brief.
Generate character exchanges and atmospheric audio with controllable dialogue timing and consistent voice direction.
Produce expressive multi-character narration, then use returned subtitle timing to align text or downstream edits.
Test dialogue-led audio in more than 20 supported language and locale options before a human linguistic and cultural review.
Automate two-minute-or-shorter assets with reference inputs, multiple output formats, subtitle data, and watermark controls.
Capabilities
Combines speech, multiple speakers, sound effects, ambience, and other scene audio in one generated result.
The launch describes prompt-level dialogue timing control in 100 ms intervals for coordinating scripted lines.
Generate from text alone, guide with up to three audio clips, select a supported or cloned speaker, or use one reference image. Audio and image references cannot be mixed in the same request.
Describe a voice in text, use an authorized reference, or combine voice direction with references to guide timbre, emotion, rhythm, and speaking style.
The API lists English, Chinese, Japanese, Korean, European and Latin American languages, and several Southeast Asian languages and locales.
The API supports an original model output of up to 120 seconds, while ByteDance also describes continuation for longer creative sequences.
Optionally return word-level timestamps and choose WAV, MP3, PCM, or Ogg Opus with configurable sample rate, speed, pitch, and loudness.
The API can add an audible rhythm marker and implicit metadata to identify AI-generated audio.
Process
Step 1
Specify speakers, exact dialogue, vocal direction, environment, effects, pacing, and the intended length instead of relying on a vague mood prompt.
Step 2
Use text only, up to three audio references, a speaker ID, or one image. Obtain permission for every voice, clip, image, and character likeness before uploading it.
Step 3
Add millisecond start times when lines must align with a visual edit, then leave room for pauses, effects, and ambience.
Step 4
Validate pronunciation, voice identity, timing, scene balance, and prompt adherence with a short sample before generating the full asset.
Step 5
Listen for artifacts, unwanted impersonation, incorrect language, clipped speech, timing drift, and unsafe content, then mix or edit the approved output in an audio editor.
Cost
BytePlus bills Seed Audio 1.0 by the model's original generated duration with second-level precision. Pay-as-you-go is $0.15 per minute, and service activation includes 60 trial minutes. Prepaid packages reduce the effective minute price but expire after one year.
60 generated minutes
A one-time usage allowance provided when the Audio 1.0 service is activated.
$0.15 per minute
Usage-based API billing with no prepaid volume commitment.
$28.50
Small prepaid package with a one-year validity period.
$27,000
High-volume annual package for production workloads.
$64,000
Larger annual package with a lower effective minute rate.
$240,000
The largest package listed on the public pricing page.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Content Creator
Choose ElevenLabs when a broader creator platform combining voice, music, sound effects, images, and video is more useful than a focused API.
Explore ElevenLabs Image & Video →Content Creator
Consider Stable Audio 3.0 for an open-weight, fully licensed audio-model family and music-oriented generation workflows.
Explore Stable Audio 3.0 →Content Creator
Use Suno when complete songs, vocals, lyrics, and music creation are the core requirement.
Explore Suno AI →Questions
Seed Audio 1.0 is ByteDance's full-scene audio-generation model for creating dialogue, speech, sound effects, ambience, and other scene audio from one prompt. It is available through BytePlus.
The BytePlus API costs $0.15 per generated minute, billed to the second using the model's original output duration. Activating the service includes 60 trial minutes, and one-year prepaid packages start at $28.50 for 200 minutes.
The current API supports up to 120 seconds of original model output in one request. ByteDance also describes continuation as part of the model's longer-form workflow.
It can create full audio scenes that include speech, effects, ambience, and other elements, but ByteDance lists finer music timing control as future work and offers Seed-Music as a separate specialist model. Use a dedicated music generator when songs or detailed composition are the main objective.
The API can use a cloned speaker or authorized reference audio to guide a voice. Only use voices and recordings you have permission to use, and review outputs for impersonation and disclosure risks.
The current BytePlus API lists more than 20 language and locale options, including English, Chinese, Japanese, Korean, French, German, Italian, Russian, several Spanish and Portuguese locales, and multiple Southeast Asian languages.
Yes. A request can use up to three audio references, a supported or cloned speaker, or one image reference. Image and audio references cannot be mixed in the same request.
Yes. The API can return utterance- and word-level subtitle timing when subtitle output is enabled.
Bottom line
Seed Audio 1.0 is a compelling API for short, dialogue-led scenes that would otherwise require separate voice, effect, and ambience passes. Its $0.15-per-minute list price is transparent, and the 60-minute trial is enough for a meaningful acceptance-rate test. The important boundary is control: dialogue timing is the mature focus today, while detailed control over effects, ambience, and music remains on ByteDance's roadmap. Use it for scene composition, not as an automatic replacement for a dedicated music model or final human audio review.
Visit Seed Audio 1.0 website ↗
Seedance 2.0 - ByteDance's frontier AI video model

Runway Dev- New developer API for Runway's media models and real-time avatars
.gif)
Nano Banana 2 Lite - Google's high-volume, cost-effective image model that generates pictures in just four seconds

Lucy 2.5 - Decart's AI that edits and restyles video mid-stream

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.