Recevez 100 crédits gratuits à l'inscription pour explorer et créer vos applications d'IAObtenir
qwen3-tts-instruct-flash

Qwen3 TTS Instruct Flash / Texte vers Parole

Commercial
ID: qwen3-tts-instruct-flash

Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments.

Prix$0.011/ 1K Characters(11 Crédits)
24 total
Synthèse vocale (TTS)
0
Sortie audioPrêt pour la synthèse. Cliquez sur « Démarrer la synthèse » pour générer la parole.
0:00 / 0:00
Qwen3 TTS Instruct Flash • Style-Controlled Speech

Qwen3 TTS Instruct FlashInstructable Multilingual Text-to-Speech

Qwen3 TTS Instruct Flash synthesizes speech with named voices, natural-language style instructions, eight language types, and optional DashScope SSE streaming — tone you direct, not just select.

Named Voices
Style Instructions
8 Language Types
SSE Streaming
optimize_instructions
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
Qwen3 TTS Instruct Flash Hero

At a Glance

Flash
Latency Tier
Qwen3 TTS
8
Languages
Auto → Japanese
Style
Instructions
≤ 1600 tokens
SSE
Streaming
X-DashScope-SSE
Capability Highlights

Instructable Speech Stack

Control tone with a director’s note — not just a voice ID.

Style Instructions

instructions sets delivery style (up to ~1600 tokens, Chinese or English). optimize_instructions can refine the directive before synthesis so first takes land closer to brief.

instructions≤ 1600 tokensoptimize_instructions
Style instructions
Voice library

Named Voice Library

Default voice Cherry, plus Serena, Ethan, Chelsie, Momo, Vivian, Moon, Maia, Kai, Nofish, Bella, Mia, Vincent, Bunny, Neil, Elias and more — a cast, not a numeric menu.

voice: CherryMultilingual CastNamed IDs

Language & Stream

language_type covers Auto, Chinese, English, German, Italian, Portuguese, Spanish, and Japanese. X-DashScope-SSE enable returns text/event-stream for progressive audio.

8 LanguagesX-DashScope-SSEAuto Detect
Language & Stream
How It Works

How It Works

From script to styled speech.

01

Write Input

Send the text to synthesize in the input field.

02

Pick Voice

Choose a named voice such as Cherry (default) or Ethan.

03

Direct Style

Add instructions for mood and pace; optionally enable optimize_instructions.

04

Select Language

Set language_type (Auto by default) and enable X-DashScope-SSE if streaming.

Qwen3 TTS Domains

Where tone must be directed, not just selected.

Product

Localized Product VO

Eight language types with style-matched delivery.

i18n
Games

Character Dialogue

Instructions for mood, pace, and persona.

Style
Education

Learning Content

Clear instructional reads with stable voices.

Edu
A11y

Accessibility Reads

On-demand narration of UI and documents.

A11y
Best Practices

Usage Tips

More reliable style control on the first pass.

Write instructions like a director

Use phrases such as “calm, slower, warm smile in the voice” — not just “happy”.

Match language_type to script

Set Chinese / English explicitly when Auto mis-detects mixed technical copy.

Keep instructions short

Stay well under the 1600-token instruction budget so the text dominates the take.

Developer Quickstart

Qwen3 TTS Instruct Flash Quickstart

Style-directed multilingual speech synthesis.

curl -X POST "https://api.powertokens.ai/v1/audio/speech" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "qwen3-tts-instruct-flash",
  "input": "欢迎收听本集产品播报,语气温暖、语速适中。",
  "voice": "Cherry",
  "speed": 1,
  "response_format": "mp3"
}' \
  --output speech.mp3

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
Alibaba / Qwen • qwen3-tts-instruct-flash
Capability
Text-to-Speech (POST /v1/audio/speech)
Input
input text to synthesize
Voice
voice default Cherry (named multilingual cast)
Instructions
style directions ≤ 1600 tokens (Chinese / English); optimize_instructions optional
Language Type
Auto, Chinese, English, German, Italian, Portuguese, Spanish, Japanese
Streaming
X-DashScope-SSE: enable returns text/event-stream
Protocol
DashScope-compatible audio speech
Billing
Usage-based via pricing matrix
API Endpoint
POST /v1/audio/speech

Détails des tarifs

La facturation réelle de ce modèle est calculée dynamiquement en fonction des paramètres spécifiques de votre requête API. Voici les combinaisons spécifiques et leurs tarifs correspondants :

Standard
Crédits11/ 1K Characters
Prix (USD)$0.011
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
MiniMax Speech 2.8 HD
audioCommercial
MiniMax Speech 2.8 HD

MiniMax Speech 2.8 HD

speech-2.8-hd

speech-2.8-hd is a high-definition AI speech synthesis model tailored for individual users. It delivers studio-grade natural voice texture with ultra-realistic pronunciation and smooth intonation. It supports rich exclusive timbres and multilingual conversion, and is capable of simulating vivid emotions like laughter and sighs. It perfectly fits daily voice dubbing, audio creation, reading narration and personal voice customization, bringing you immersive and high-quality voice experience.

Text to Music
MiniMax Speech 2.8 Turbo
audioCommercial
MiniMax Speech 2.8 Turbo

MiniMax Speech 2.8 Turbo

speech-2.8-turbo

speech-2.8-turbo is a lightweight and ultra-fast AI speech synthesis model for all users. It features instant response, efficient generation and stable audio output. With natural and smooth timbre performance, it supports multilingual conversion and basic emotional intonation adjustment. Optimized for low-latency scenarios such as daily narration, short video dubbing and real-time voice interaction, it balances speed, quality and ease of use, delivering a fluent and convenient voice creation exp

Text to Music
MiniMax Speech 2.6 HD
audioCommercial
MiniMax Speech 2.6 HD

MiniMax Speech 2.6 HD

speech-2.6-hd

speech-2.6-hd is a high-definition AI voice synthesis model designed for general users. It delivers lifelike, studio-level vocal quality with natural pronunciation, smooth rhythm and rich emotional expression. It supports multiple languages and diverse premium voice tones, enabling vivid voice dubbing, audiobook narration and personalized voice creation. With stable sound quality and authentic intonation, it perfectly fits daily entertainment, content creation and daily voice playback needs, bri

Text to Music
MiniMax Speech 2.6 Turbo
audioCommercial
MiniMax Speech 2.6 Turbo

MiniMax Speech 2.6 Turbo

speech-2.6-turbo

A lightweight, ultra-fast AI speech synthesis model with premium sound quality and sub-250ms low latency. Supports 40+ languages, 7 emotions, and 300+ curated voices for real-time interaction, short video dubbing, and daily narration. Delivers natural, smooth audio with high cost-performance.

Text to Music

Frequently Asked Questions

Everything you need to know before integrating this model.

Use instructions (Chinese or English, up to ~1600 tokens) to describe mood and pace. Optionally enable optimize_instructions.

Start Building with Qwen3 TTS Instruct Flash Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.