The Rundown AI homepage

Independent tool overview

Qwen3-TTS CustomVoice 1.7B at a glance

Qwen3-TTS CustomVoice 1.7B is an Apache-licensed text-to-speech model that speaks ten languages through nine built-in voices, with streaming output and natural-language control over delivery.

Visit the official Qwen3-TTS CustomVoice 1.7B site ↗
Qwen3-TTS CustomVoice 1.7B product preview
Product type
Open-weight multilingual text-to-speech model
Model size
1.7 billion parameters
Built-in speakers
Nine preset voices
Languages
Ten documented languages
Output modes
Streaming and non-streaming
License
Apache 2.0
Pricing checked
August 31, 2026

Overview

What Qwen3-TTS CustomVoice 1.7B is

Qwen3-TTS CustomVoice 1.7B is the preset-speaker member of Qwen's open speech-generation family. It turns text into speech using nine named voices and can follow instructions about emotion, pace, tone and prosody.

The model supports Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Qwen recommends using each speaker's native language for the best quality, although every listed speaker can generate any supported language.

Despite the original directory description, this CustomVoice checkpoint does not clone a person's voice from a three-second recording. Rapid voice cloning belongs to the separate Qwen3-TTS Base checkpoints; VoiceDesign is another separate model. That distinction matters for both product selection and consent controls.

Use cases

Who Qwen3-TTS CustomVoice 1.7B is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Self-hosted narration

Teams generating multilingual product narration, explainers or accessibility audio with an open model and a fixed voice roster.

Expressive scripted speech

Dialogue and character lines that benefit from instructions for emotion, rhythm, speaking rate and delivery.

Low-latency prototypes

Developers testing streaming speech for assistants or interactive experiences before selecting production infrastructure.

Capabilities

Core Qwen3-TTS CustomVoice 1.7B features

1

Nine preset timbres

The official roster includes Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna and Sohee.

2

Ten-language synthesis

Generates Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian.

3

Instruction-based delivery

A free-text instruction can guide emotion, tone, rate and prosody for the 1.7B checkpoint.

4

Streaming architecture

The family supports streaming and non-streaming generation; Qwen reports end-to-end first-packet latency as low as 97 ms under its test conditions.

5

Batch generation

The Python interface accepts lists of text, language, speaker and instruction values for multiple outputs.

6

Local Python package and demo

The qwen-tts package loads checkpoints directly and includes a local web demo command.

7

Open model files

Weights are available through Hugging Face and ModelScope under Apache 2.0, with vLLM-Omni offline-inference examples.

Process

How the Qwen3-TTS CustomVoice 1.7B workflow works

  1. Step 1

    Choose the correct Qwen3-TTS checkpoint

    Use CustomVoice for the nine preset speakers, VoiceDesign for a described synthetic persona, or Base for consented reference-audio cloning.

  2. Step 2

    Install in an isolated environment

    Use the current qwen-tts package and a compatible GPU stack; Qwen recommends FlashAttention 2 to reduce memory use where supported.

  3. Step 3

    Select language, speaker and direction

    Set the known target language explicitly, choose a speaker and write a concise performance instruction without putting secrets in the text.

  4. Step 4

    Generate short review batches

    Test pronunciation, names, numbers, code-switching, emotion, noise, latency and consistency across the actual scripts and target devices.

  5. Step 5

    Approve and disclose the audio

    Have a fluent reviewer listen to the full output, correct errors, document the synthetic voice and obtain the rights needed for script, music and distribution.

Cost

Qwen3-TTS CustomVoice 1.7B pricing and free plan

The CustomVoice checkpoint is free under Apache 2.0, but self-hosting requires suitable compute. Alibaba Cloud also sells related managed Qwen3-TTS speech APIs; those model IDs are hosted services and should not be assumed to be the identical 1.7B checkpoint.

Open model

Free under Apache 2.0

Download and operate the 1.7B CustomVoice weights.

  • No model license fee
  • Hardware, storage and engineering are not included
  • Use of voices and generated content remains subject to applicable rights and law

Self-hosted inference

Compute-dependent

Run locally or on rented accelerators using qwen-tts or a compatible runtime.

  • Cost varies by GPU, traffic and latency target
  • Streaming production requires your own serving and monitoring
  • Benchmark quality and real-time factor on target hardware

Alibaba Cloud Qwen3-TTS Instruct API

$0.115 per 10,000 characters

International Singapore list price for the related qwen3-tts-instruct-flash hosted model.

  • Output is free under the listed billing rule
  • 110,000-character introductory quota under documented conditions
  • Real-time Instruct pricing is $0.143 per 10,000 characters
  • Hosted model IDs are not labeled as the open 1.7B checkpoint

Pricing checked . Check current pricing at the source ↗

Assessment

Qwen3-TTS CustomVoice 1.7B strengths and limitations

Where it stands out

  • Open weights and Apache 2.0 licensing support private, controlled deployments.
  • Ten documented languages cover several major narration markets.
  • Nine distinct preset speakers simplify repeatable voice selection.
  • Instruction control adds emotion and pacing without training a custom voice.
  • Streaming and batch interfaces cover interactive and offline workflows.
  • The official package, demo and model cards provide a relatively direct setup path.

What to consider

  • CustomVoice is not the three-second voice-cloning model; reference-audio cloning requires a separate Base checkpoint.
  • Only nine preset speakers are included, and voice identity cannot be freely specified the way VoiceDesign allows.
  • Qwen recommends a speaker's native language for best quality, so cross-language output may be less natural or consistent.
  • The reported 97 ms latency is a vendor result under particular conditions, not a promise for every device or deployment.
  • Names, numbers, abbreviations, mixed-language text and domain terminology can be mispronounced.
  • Emotional instructions can produce inconsistent intensity, pacing or unintended performance choices.
  • Synthetic speech can be misused for impersonation, fraud or deceptive media; applications need authentication, consent, disclosure and abuse monitoring.
  • High-impact uses such as emergency, medical, legal, financial or identity verification require human review and must not rely on generated speech alone.

Compare

Qwen3-TTS CustomVoice 1.7B alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Marketing

ElevenLabs

A managed commercial speech platform with voice libraries, cloning and production tooling for teams that do not want to self-host.

Explore ElevenLabs

Content Creator

VibeVoice

An open Microsoft speech model oriented toward long-form and multi-speaker generation.

Explore VibeVoice

Questions

Qwen3-TTS CustomVoice 1.7B FAQs

What is Qwen3-TTS CustomVoice 1.7B?

It is an open Qwen text-to-speech checkpoint with nine preset speakers, ten languages, streaming generation and instruction control over delivery.

Can this model clone a voice from three seconds of audio?

No. The three-second rapid-cloning capability belongs to Qwen3-TTS Base checkpoints. CustomVoice uses a fixed list of preset speakers.

Which languages does it support?

Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian. Qwen recommends each preset speaker's native language for the strongest quality.

Which voices are included?

Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna and Sohee. They cover Chinese and regional Chinese profiles plus native English, Japanese and Korean voices.

Is Qwen3-TTS CustomVoice open source?

The model weights and repository are published under Apache 2.0. You still pay for infrastructure and must comply with content, voice, privacy and distribution rights.

Can it run locally?

Yes, on compatible hardware. Qwen documents its qwen-tts Python package and local demo, recommends FlashAttention 2 where supported, and provides vLLM-Omni offline examples.

How much does it cost?

The open checkpoint has no model license fee. Self-hosting costs depend on hardware and traffic. A related Alibaba Cloud Instruct API is listed at $0.115 per 10,000 characters in Singapore, but it is not identified as the identical 1.7B checkpoint.

What safeguards should a speech app add?

Obtain rights to scripts and voices, disclose synthetic audio, block impersonation and fraud, protect logs and credentials, watermark or retain provenance where appropriate, and require human approval for sensitive messages.

Bottom line

Our Qwen3-TTS CustomVoice 1.7B verdict

Qwen3-TTS CustomVoice 1.7B is a practical open option when nine stable preset speakers and multilingual instruction control are enough. Its biggest documentation trap is the product boundary: it does not perform the family's three-second voice cloning, so teams needing a copied or newly designed identity must choose another checkpoint and apply stricter consent controls.

Visit Qwen3-TTS CustomVoice 1.7B website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.