
Qwen3 TTS Instruct Flash / Text to Speech
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments.
Qwen3 TTS Instruct FlashInstructable Multilingual Text-to-Speech
Qwen3 TTS Instruct Flash synthesizes speech with named voices, natural-language style instructions, eight language types, and optional DashScope SSE streaming — tone you direct, not just select.

At a Glance
Instructable Speech Stack
Control tone with a director’s note — not just a voice ID.
Style Instructions
instructions sets delivery style (up to ~1600 tokens, Chinese or English). optimize_instructions can refine the directive before synthesis so first takes land closer to brief.


Named Voice Library
Default voice Cherry, plus Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish, Bella, Mia, Vincent, Bunny, Neil, Elias and more — a cast, not a numeric menu.
Language & Stream
language_type covers Auto, Chinese, English, German, Italian, Portuguese, Spanish, and Japanese. X-DashScope-SSE enable returns text/event-stream for progressive audio.

How It Works
From script to styled speech.
Write Input
Send the text to synthesize in the input field.
Pick Voice
Choose a named voice such as Cherry (default) or Ethan.
Direct Style
Add instructions for mood and pace; optionally enable optimize_instructions.
Select Language
Set language_type (Auto by default) and enable X-DashScope-SSE if streaming.
Qwen3 TTS Domains
Where tone must be directed, not just selected.
Localized Product VO
Eight language types with style-matched delivery.
Character Dialogue
Instructions for mood, pace, and persona.
Learning Content
Clear instructional reads with stable voices.
Accessibility Reads
On-demand narration of UI and documents.
Usage Tips
More reliable style control on the first pass.
Use phrases such as “calm, slower, warm smile in the voice” — not just “happy”.
Set Chinese / English explicitly when Auto mis-detects mixed technical copy.
Stay well under the 1600-token instruction budget so the text dominates the take.
Qwen3 TTS Instruct Flash Quickstart
Style-directed multilingual speech synthesis.
curl -X POST "https://api.powertokens.ai/v1/audio/speech" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-tts-instruct-flash",
"input": "欢迎收听本集产品播报,语气温暖、语速适中。",
"voice": "Cherry",
"speed": 1,
"response_format": "mp3"
}' \
--output speech.mp3Technical Specifications
Confirmed parameters and runtime execution protocols.
Pricing Details
The actual billing for this model is dynamically calculated based on the specific parameters passed in your API request. Below are the specific combinations and their corresponding pricing:
| Modality | Credits | Price (USD) |
|---|---|---|
| Standard | 11/ 1K Characters | $0.011 |
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


MiniMax Speech 2.8 HD
speech-2.8-hd
speech-2.8-hd is a high-definition AI speech synthesis model tailored for individual users. It delivers studio-grade natural voice texture with ultra-realistic pronunciation and smooth intonation. It supports rich exclusive timbres and multilingual conversion, and is capable of simulating vivid emotions like laughter and sighs. It perfectly fits daily voice dubbing, audio creation, reading narration and personal voice customization, bringing you immersive and high-quality voice experience.


MiniMax Speech 2.8 Turbo
speech-2.8-turbo
speech-2.8-turbo is a lightweight and ultra-fast AI speech synthesis model for all users. It features instant response, efficient generation and stable audio output. With natural and smooth timbre performance, it supports multilingual conversion and basic emotional intonation adjustment. Optimized for low-latency scenarios such as daily narration, short video dubbing and real-time voice interaction, it balances speed, quality and ease of use, delivering a fluent and convenient voice creation exp


MiniMax Speech 2.6 HD
speech-2.6-hd
speech-2.6-hd is a high-definition AI voice synthesis model designed for general users. It delivers lifelike, studio-level vocal quality with natural pronunciation, smooth rhythm and rich emotional expression. It supports multiple languages and diverse premium voice tones, enabling vivid voice dubbing, audiobook narration and personalized voice creation. With stable sound quality and authentic intonation, it perfectly fits daily entertainment, content creation and daily voice playback needs, bri


MiniMax Speech 2.6 Turbo
speech-2.6-turbo
A lightweight, ultra-fast AI speech synthesis model with premium sound quality and sub-250ms low latency. Supports 40+ languages, 7 emotions, and 300+ curated voices for real-time interaction, short video dubbing, and daily narration. Delivers natural, smooth audio with high cost-performance.
Frequently Asked Questions
Everything you need to know before integrating this model.
Start Building with Qwen3 TTS Instruct Flash Today
Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.