
MiniMax Speech-02 HD / Texto para Voz
A high-definition AI speech model with outstanding prosody, stability, and industry-leading voice cloning similarity. Generates studio-grade, ultra-natural voices with rich emotional expression. Ideal for audiobooks, premium voiceovers, and professional content creation in 40+ languages.
Speech 02 HDHigh-fidelity studio TTS
Speech 02 HD is MiniMax 02-series speech synthesis with selectable Mandarin voices, speed control, and mp3 / wav / pcm output — tuned for high-fidelity studio narration. Prefer yujie or chengshu for premium brand VO.

At a Glance
Speech 02 HD Speech Architecture
A complete TTS control surface — voice identity, prosody, container format, and delivery mode.
HD Voice Fidelity
Prioritizes natural prosody, breath, and studio-grade clarity for narration, IVR, and brand voice assets that ship as finished media.


Voice Library & Persona
Default voice male-qn-qingse, with male and female Mandarin personas including shaonv, yujie, chengshu, tianmei, jingying, badao, and daxuesheng. Prefer yujie or chengshu for premium brand VO.
Pace & Container Controls
speed maps to voice_setting.speed (default 1.0, range 0.5–2.0). response_format supports mp3, wav, and pcm so post pipelines can pick the right container.


Delivery & Streaming
metadata.output_format switches between url and hex delivery (url default). stream_format enables upstream streaming when any non-empty value is sent.
End-to-End Speech Pipeline
From script to finished, mix-ready audio in five deliberate steps.
Write Input Script
Send up to 10,000 characters of punctuated plain text. Written rhythm drives pauses and emphasis.
Cast the Voice
Choose a voice_id such as male-qn-qingse, female-yujie, or female-tianmei to match brand persona.
Direct Pace
Set speed between 0.5 and 2.0. Slow down for legal or medical copy; speed up for UI microcopy.
Pick Container
Select mp3 for direct delivery, wav/pcm for mix and master. Use url or hex via metadata.output_format.
Stream or File
Set stream_format when you need progressive chunks; otherwise download the finished asset.
Speech 02 HD Production Domains
Where this tier of MiniMax Speech ships every week.
Brand Narration
Premium VO for product films, explainers, and launch films.
Audiobook Drafts
Long-form input up to 10,000 characters per synthesis call.
Podcast Intros
Consistent host identity across weekly episode openers.
Course Modules
Clear instructional reads with stable pacing and diction.
Voice Direction Notes
Studio craft that lifts first-pass quality more than any single parameter.
Use commas, periods, and line breaks to control pauses — the model follows written rhythm closely.
Pick shaonv / tianmei for bright product UI; yujie / chengshu for premium brand VO; daxuesheng for campus tone.
Use wav or pcm when you will mix or master; mp3 for direct CDN delivery to end users.
Write “twenty twenty-four” instead of “2024” when you need spoken years or codes to land cleanly.
Speech 02 HD Quickstart
Synthesize speech with explicit voice, pace, and format controls.
curl -X POST "https://api.powertokens.ai/v1/audio/speech" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "speech-02-hd",
"input": "欢迎使用 MiniMax 语音合成,这是一段示例播报。",
"voice": "male-qn-qingse",
"speed": 1,
"response_format": "mp3"
}' \
--output speech.mp3Technical Specifications
Confirmed parameters and runtime execution protocols.
Detalhes de preços
A cobrança real deste modelo é calculada dinamicamente com base nos parâmetros específicos da sua solicitação de API. Abaixo estão as combinações específicas e seus preços correspondentes:
| Modalidade | Créditos | Preço (USD) |
|---|---|---|
| Standard | 95/ 1K Characters | $0.095 |
Models from the Same Channel
Explore complementary models and alternative versions from the same provider channel.


MiniMax Speech 2.8 HD
speech-2.8-hd
speech-2.8-hd is a high-definition AI speech synthesis model tailored for individual users. It delivers studio-grade natural voice texture with ultra-realistic pronunciation and smooth intonation. It supports rich exclusive timbres and multilingual conversion, and is capable of simulating vivid emotions like laughter and sighs. It perfectly fits daily voice dubbing, audio creation, reading narration and personal voice customization, bringing you immersive and high-quality voice experience.


MiniMax Speech 2.8 Turbo
speech-2.8-turbo
speech-2.8-turbo is a lightweight and ultra-fast AI speech synthesis model for all users. It features instant response, efficient generation and stable audio output. With natural and smooth timbre performance, it supports multilingual conversion and basic emotional intonation adjustment. Optimized for low-latency scenarios such as daily narration, short video dubbing and real-time voice interaction, it balances speed, quality and ease of use, delivering a fluent and convenient voice creation exp


MiniMax Speech 2.6 HD
speech-2.6-hd
speech-2.6-hd is a high-definition AI voice synthesis model designed for general users. It delivers lifelike, studio-level vocal quality with natural pronunciation, smooth rhythm and rich emotional expression. It supports multiple languages and diverse premium voice tones, enabling vivid voice dubbing, audiobook narration and personalized voice creation. With stable sound quality and authentic intonation, it perfectly fits daily entertainment, content creation and daily voice playback needs, bri


MiniMax Speech 2.6 Turbo
speech-2.6-turbo
A lightweight, ultra-fast AI speech synthesis model with premium sound quality and sub-250ms low latency. Supports 40+ languages, 7 emotions, and 300+ curated voices for real-time interaction, short video dubbing, and daily narration. Delivers natural, smooth audio with high cost-performance.
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


Qwen3 TTS Instruct Flash
qwen3-tts-instruct-flash
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments.


Vidu Audio 1.0
audio1.0
Vidu's text-to-audio generation model (model ID: audio1.0). Generates sound effects and background music from text prompts. Duration range: 2–10 seconds. Supports random seed configuration for reproducible outputs
Frequently Asked Questions
Everything you need to know before integrating this model.
Start Building with Speech 02 HD Today
Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.