The Rundown AI homepage

Independent tool overview

Fish Audio S1 at a glance

Fish Audio S1 is a previous-generation 4-billion-parameter text-to-speech model with voice cloning, 13-language speech generation, and 64-plus parenthetical emotion and delivery controls. It remains available in the Fish Audio API for existing integrations, but Fish Audio recommends S2.1-Pro for new production projects.

Visit the official Fish Audio S1 site ↗
Fish Audio S1 product preview
Current position
Previous model; supported for existing integrations
Recommended successor
S2.1-Pro
Model size
4 billion parameters
Languages
13
Emotion controls
64+ parenthetical expressions
API price
$15 per million UTF-8 input bytes
Multi-speaker dialogue
No; use S2.1-Pro
Last reviewed
August 29, 2026

Overview

What Fish Audio S1 is

S1 turns text into expressive single-speaker audio using a selected Fish Audio voice or a user-created voice clone. Its distinctive control system uses parenthetical cues such as (whispering), (excited), (sighing), or (sarcastic) inside the script to influence delivery.

Developers can call S1 with the model header s1 through Fish Audio’s text-to-speech API and request MP3, WAV/PCM, or Opus output. The broader platform also includes a browser playground, voice creation, model visibility controls, real-time streaming, JavaScript and Python SDKs, and pay-as-you-go API billing.

S1 is now a compatibility choice rather than the default for new builds. S2.1-Pro adds 83 languages, multi-speaker dialogue, natural-language bracket controls, and improved quality, latency, and throughput. Fish Audio also offers the same newer model through s2.1-pro-free for development and testing under fair-use limits, without production TTFA or DPA guarantees.

Use cases

Who Fish Audio S1 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Existing S1 applications

Keep a production integration stable while testing output and cost against the newer S2.1-Pro model.

Expressive single-speaker narration

Generate audiobooks, character lines, explainers, podcast segments, and other scripts that benefit from explicit emotion and delivery cues.

Multilingual voice clones

Use a permitted reference voice across S1’s supported languages after testing pronunciation, accent, and identity consistency.

Developer-controlled TTS

Integrate generated speech through REST, WebSocket streaming, Python, or JavaScript with configurable format, prosody, and sampling parameters.

Capabilities

Core Fish Audio S1 features

1

64+ expression controls

S1 interprets parenthetical emotion, tone, and sound cues including excited, nervous, whispering, laughing, sighing, and many others.

2

13 supported languages

S1 supports English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, and Portuguese.

3

Voice cloning

Users can create a voice model from permitted reference audio and use its reference ID for repeated synthesis.

4

Zero-shot references

The API can also accept reference audio and transcript data directly for a synthesis request instead of relying only on a pre-created voice.

5

Multiple output formats

The API supports mono WAV/PCM, MP3, and Opus with documented sample-rate and bitrate choices.

6

Prosody controls

Developers can adjust speed, volume, normalization, temperature, top-p, chunking, and other synthesis parameters.

7

Streaming and SDKs

Fish Audio documents real-time WebSocket generation plus Python and JavaScript clients for application integration.

8

Voice visibility settings

Created models can be public, unlisted, or private, allowing teams to separate discoverable characters from controlled proprietary voices.

9

Web creator plans

The Fish Audio web app provides monthly credits, cloning, voice slots, commercial-use terms, and team features separate from API usage.

Process

How the Fish Audio S1 workflow works

  1. Step 1

    Confirm model fit

    For a new project, benchmark S1 only if compatibility or its parenthetical control behavior matters; otherwise start with S2.1-Pro or the free development version.

  2. Step 2

    Secure voice consent

    Use your own voice or obtain written permission that covers cloning, intended uses, languages, duration, disclosure, and revocation.

  3. Step 3

    Record clean reference audio

    Capture at least 10 seconds of isolated speech without music, other speakers, reverberation, or background noise; longer clean clips can improve consistency.

  4. Step 4

    Set the correct visibility

    Choose private for controlled voices. The create-model API defaults to public, which exposes the model in discovery unless changed.

  5. Step 5

    Prepare and mark up the script

    Normalize names, numbers, abbreviations, and pronunciation, then add S1 parenthetical cues sparingly at the moments that need direction.

  6. Step 6

    Generate short test segments

    Evaluate pronunciation, emotional delivery, pauses, identity similarity, artifacts, and transitions before synthesizing long material.

  7. Step 7

    Add disclosure and safeguards

    Label synthetic audio where appropriate and block impersonation, fraud, deceptive calls, or other unapproved uses.

  8. Step 8

    Measure the migration path

    Compare S1 with S2.1-Pro on quality, languages, multi-speaker needs, latency, throughput, API cost, and required code changes.

Cost

Fish Audio S1 pricing and free plan

Fish Audio sells creator subscriptions for the web app and separate usage-based API access. S1 API synthesis is $15 per million UTF-8 input bytes; Fish Audio estimates that amount as roughly 180,000 English words or about 12 hours of speech, but actual duration varies.

Free web plan

$0

Entry plan for limited monthly voice generation.

  • 8,000 credits monthly
  • Up to 7 minutes of generation
  • 500 characters per generation
  • 3 public voice slots
  • No credit card required

Plus

$15/month or $66/year

Creator plan with larger jobs, private voice slots, Voice Design, and priority generation.

  • Annual promotional equivalent shown as $5.50/month
  • 250,000 credits monthly
  • Up to 200 minutes
  • 15,000 characters per generation
  • 10 private voice slots

Pro

$100/month or $450/year

Power-user and business plan with team seats and larger production limits.

  • Annual promotional equivalent shown as $37.50/month
  • 2,000,000 credits monthly
  • Up to 1,620 minutes
  • 3 team seats
  • 30,000 characters per generation

Max

$999/month or $8,988/year

High-volume creator plan with more credits, minutes, seats, and professional voice slots.

  • Annual equivalent shown as $749/month
  • 25,000,000 credits monthly
  • Up to 6,250 minutes
  • 10 team seats
  • 15 professional voice slots

S1 API

$15 / million UTF-8 bytes

Pay-as-you-go synthesis using the s1 model header.

  • No API subscription fee or monthly minimum documented
  • Approximately 180,000 English words per million bytes
  • Five concurrent requests below $100 prepaid
  • Higher concurrency unlocks at prepaid thresholds

Enterprise

Custom

Annual volume arrangement for organizations with compliance and deployment requirements.

  • Organization-level controls
  • Zero data retention advertised
  • On-premise deployment
  • SOC 2 compliance advertised
  • Custom SSO listed as coming soon

Pricing checked . Check current pricing at the source ↗

Assessment

Fish Audio S1 strengths and limitations

Where it stands out

  • Large library of explicit emotion, tone, and nonverbal delivery cues
  • Supports 13 languages with voice cloning and cross-language generation workflows
  • Available through a documented REST API, WebSocket streaming, and two official SDKs
  • S1 remains supported, reducing forced migration risk for existing integrations
  • Usage-based API price is published and does not require a monthly minimum
  • Web plans cover a wide range from free experiments to high-volume team production
  • Private and unlisted voice model settings are available when configured deliberately

What to consider

  • Fish Audio classifies S1 as a previous model and recommends S2.1-Pro for all new production projects.
  • S1 supports 13 languages versus 83 for S2.1-Pro and does not support the newer model’s multi-speaker dialogue.
  • Its parenthetical cue syntax differs from the natural-language bracket controls in S2.1, creating migration work for marked-up scripts.
  • The published word-error, character-error, speed, and arena-ranking figures are vendor-reported benchmarks and may not predict a specific voice, language, or script.
  • Generated speech can mispronounce names, skip or repeat content, add artifacts, or produce an emotion different from the requested cue.
  • Voice cloning can enable impersonation, fraud, harassment, and deceptive media; written consent and use-specific controls are essential.
  • The create-model API defaults visibility to public, so developers must explicitly select private when a voice should not appear in discovery.
  • Web-plan minutes are estimates based on credit consumption and reset monthly without rollover.
  • Annual plan prices currently include promotional discounts and may change at renewal.
  • API cost is based on UTF-8 bytes rather than finished minutes, so non-ASCII languages and output duration can produce different economics.
  • S1 is single-speaker per request; dialogue requires separate generation and editing or migration to S2.1-Pro.
  • Commercial rights, voice ownership, privacy, disclosure, labor agreements, publicity rights, and local synthetic-media law still apply even when the platform allows generation.

Compare

Fish Audio S1 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Miscellaneous

TADA by Hume AI

TADA by Hume AI is an open-source alternative designed around tightly aligned text and speech generation.

Explore TADA by Hume AI

Business Operations

Speechify

Speechify is better suited to users who primarily want to listen to documents and web content through a consumer reading experience.

Explore Speechify

Content Creator

ElevenLabs Text To Speech GPT

The ElevenLabs custom GPT offers a simpler ChatGPT-based path into speech generation, with different platform and usage constraints.

Explore ElevenLabs Text To Speech GPT

Questions

Fish Audio S1 FAQs

What is Fish Audio S1?

S1 is Fish Audio’s previous-generation 4-billion-parameter text-to-speech model for expressive single-speaker audio, voice cloning, and 13-language generation.

Is Fish Audio S1 still available?

Yes. Fish Audio keeps S1 available for existing integrations, but recommends S2.1-Pro for new production projects.

How much does the Fish Audio S1 API cost?

S1 costs $15 per million UTF-8 input bytes. Fish Audio estimates that as roughly 180,000 English words or about 12 hours of speech, though real duration varies.

What is the difference between S1 and S2.1-Pro?

S2.1-Pro is the recommended production model with 83 languages, multi-speaker dialogue, natural-language bracket controls, and improved quality, latency, and throughput. S1 has 13 languages and parenthetical expression controls.

Is there a free Fish Audio model for developers?

Yes. Fish Audio documents s2.1-pro-free as the same newer model available at $0 under fair-use limits for development and testing, without production TTFA or DPA guarantees.

How many languages does S1 support?

S1 supports 13: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, and Portuguese.

Can S1 clone a voice?

Yes. Fish Audio supports created voice models and zero-shot reference audio. Only clone voices you own or have explicit written permission to use.

Can S1 generate multiple speakers in one request?

No. Fish Audio documents multi-speaker dialogue as an S2.1-Pro and S2-Pro capability, not an S1 capability.

Are Fish Audio voice models private?

They can be, but privacy is configurable. The create-model API defaults to public; choose private explicitly for restricted voices.

Bottom line

Our Fish Audio S1 verdict

Fish Audio S1 remains a capable expressive TTS option for teams with an existing integration or scripts built around its parenthetical emotion controls. New projects should usually start with S2.1-Pro—or s2.1-pro-free for development—because of the much broader language support, multi-speaker generation, and current production focus. Whichever model you choose, voice consent, private-by-default configuration, disclosure, and output QA matter as much as audio quality.

Visit Fish Audio S1 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.