Existing S1 applications
Keep a production integration stable while testing output and cost against the newer S2.1-Pro model.
Independent tool overview
Fish Audio S1 is a previous-generation 4-billion-parameter text-to-speech model with voice cloning, 13-language speech generation, and 64-plus parenthetical emotion and delivery controls. It remains available in the Fish Audio API for existing integrations, but Fish Audio recommends S2.1-Pro for new production projects.
Visit the official Fish Audio S1 site ↗
Overview
S1 turns text into expressive single-speaker audio using a selected Fish Audio voice or a user-created voice clone. Its distinctive control system uses parenthetical cues such as (whispering), (excited), (sighing), or (sarcastic) inside the script to influence delivery.
Developers can call S1 with the model header s1 through Fish Audio’s text-to-speech API and request MP3, WAV/PCM, or Opus output. The broader platform also includes a browser playground, voice creation, model visibility controls, real-time streaming, JavaScript and Python SDKs, and pay-as-you-go API billing.
S1 is now a compatibility choice rather than the default for new builds. S2.1-Pro adds 83 languages, multi-speaker dialogue, natural-language bracket controls, and improved quality, latency, and throughput. Fish Audio also offers the same newer model through s2.1-pro-free for development and testing under fair-use limits, without production TTFA or DPA guarantees.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Keep a production integration stable while testing output and cost against the newer S2.1-Pro model.
Generate audiobooks, character lines, explainers, podcast segments, and other scripts that benefit from explicit emotion and delivery cues.
Use a permitted reference voice across S1’s supported languages after testing pronunciation, accent, and identity consistency.
Integrate generated speech through REST, WebSocket streaming, Python, or JavaScript with configurable format, prosody, and sampling parameters.
Capabilities
S1 interprets parenthetical emotion, tone, and sound cues including excited, nervous, whispering, laughing, sighing, and many others.
S1 supports English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, and Portuguese.
Users can create a voice model from permitted reference audio and use its reference ID for repeated synthesis.
The API can also accept reference audio and transcript data directly for a synthesis request instead of relying only on a pre-created voice.
The API supports mono WAV/PCM, MP3, and Opus with documented sample-rate and bitrate choices.
Developers can adjust speed, volume, normalization, temperature, top-p, chunking, and other synthesis parameters.
Fish Audio documents real-time WebSocket generation plus Python and JavaScript clients for application integration.
Created models can be public, unlisted, or private, allowing teams to separate discoverable characters from controlled proprietary voices.
The Fish Audio web app provides monthly credits, cloning, voice slots, commercial-use terms, and team features separate from API usage.
Process
Step 1
For a new project, benchmark S1 only if compatibility or its parenthetical control behavior matters; otherwise start with S2.1-Pro or the free development version.
Step 2
Use your own voice or obtain written permission that covers cloning, intended uses, languages, duration, disclosure, and revocation.
Step 3
Capture at least 10 seconds of isolated speech without music, other speakers, reverberation, or background noise; longer clean clips can improve consistency.
Step 4
Choose private for controlled voices. The create-model API defaults to public, which exposes the model in discovery unless changed.
Step 5
Normalize names, numbers, abbreviations, and pronunciation, then add S1 parenthetical cues sparingly at the moments that need direction.
Step 6
Evaluate pronunciation, emotional delivery, pauses, identity similarity, artifacts, and transitions before synthesizing long material.
Step 7
Label synthetic audio where appropriate and block impersonation, fraud, deceptive calls, or other unapproved uses.
Step 8
Compare S1 with S2.1-Pro on quality, languages, multi-speaker needs, latency, throughput, API cost, and required code changes.
Cost
Fish Audio sells creator subscriptions for the web app and separate usage-based API access. S1 API synthesis is $15 per million UTF-8 input bytes; Fish Audio estimates that amount as roughly 180,000 English words or about 12 hours of speech, but actual duration varies.
$0
Entry plan for limited monthly voice generation.
$15/month or $66/year
Creator plan with larger jobs, private voice slots, Voice Design, and priority generation.
$100/month or $450/year
Power-user and business plan with team seats and larger production limits.
$999/month or $8,988/year
High-volume creator plan with more credits, minutes, seats, and professional voice slots.
$15 / million UTF-8 bytes
Pay-as-you-go synthesis using the s1 model header.
Custom
Annual volume arrangement for organizations with compliance and deployment requirements.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Miscellaneous
TADA by Hume AI is an open-source alternative designed around tightly aligned text and speech generation.
Explore TADA by Hume AI →Business Operations
Speechify is better suited to users who primarily want to listen to documents and web content through a consumer reading experience.
Explore Speechify →Content Creator
The ElevenLabs custom GPT offers a simpler ChatGPT-based path into speech generation, with different platform and usage constraints.
Explore ElevenLabs Text To Speech GPT →Questions
S1 is Fish Audio’s previous-generation 4-billion-parameter text-to-speech model for expressive single-speaker audio, voice cloning, and 13-language generation.
Yes. Fish Audio keeps S1 available for existing integrations, but recommends S2.1-Pro for new production projects.
S1 costs $15 per million UTF-8 input bytes. Fish Audio estimates that as roughly 180,000 English words or about 12 hours of speech, though real duration varies.
S2.1-Pro is the recommended production model with 83 languages, multi-speaker dialogue, natural-language bracket controls, and improved quality, latency, and throughput. S1 has 13 languages and parenthetical expression controls.
Yes. Fish Audio documents s2.1-pro-free as the same newer model available at $0 under fair-use limits for development and testing, without production TTFA or DPA guarantees.
S1 supports 13: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, and Portuguese.
Yes. Fish Audio supports created voice models and zero-shot reference audio. Only clone voices you own or have explicit written permission to use.
No. Fish Audio documents multi-speaker dialogue as an S2.1-Pro and S2-Pro capability, not an S1 capability.
They can be, but privacy is configurable. The create-model API defaults to public; choose private explicitly for restricted voices.
Bottom line
Fish Audio S1 remains a capable expressive TTS option for teams with an existing integration or scripts built around its parenthetical emotion controls. New projects should usually start with S2.1-Pro—or s2.1-pro-free for development—because of the much broader language support, multi-speaker generation, and current production focus. Whichever model you choose, voice consent, private-by-default configuration, disclosure, and output QA matter as much as audio quality.
Visit Fish Audio S1 website ↗
Hunyuan-Vision-1.5-Thinking - Tencent's most advanced vision-language model

PokeeResearch 7B- Pokee AI's SOTA, open-source deep research agent

OpenAI's voice model that listens while talking, now for developers

HunyuanWorld-Mirror - Generate 3D worlds from text, image, or video inputs

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.