Recevez 100 crédits gratuits à l'inscription pour explorer et créer vos applications d'IAObtenir
speech-02-turbo

MiniMax Speech-02 Turbo / Texte vers Parole

Commercial
ID: speech-02-turbo

An efficient turbo AI speech model with enhanced small-language support and reliable performance. Boasts low latency, fast generation, and natural intonation. Supports 40+ languages and basic emotional adjustment, perfect for real-time apps, chatbots, and cost-effective large-scale use.

Prix$0.057/ 1K Characters(57 Crédits)
327 total
1
Synthèse vocale (TTS)
0 / 10,000
Sortie audioPrêt pour la synthèse. Cliquez sur « Démarrer la synthèse » pour générer la parole.
0:00 / 0:00
Speech 02 Turbo Hero
MiniMax Speech 02 Turbo • Text-to-Speech

Speech 02 TurboLow-latency production TTS

Speech 02 Turbo is MiniMax 02-series speech synthesis with selectable Mandarin voices, speed control, and mp3 / wav / pcm output — tuned for low-latency interactive voice. Prefer jingying or qingse for crisp assistant tone.

Turbo Latency
8 Core Voice IDs
Speed 0.5–2.0
mp3 / wav / pcm
url / hex Delivery
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA

At a Glance

02
Series
MiniMax Speech 02
Turbo
Tier
Low latency
10K
Input Chars
Per request
3
Formats
mp3 · wav · pcm
0.5–2.0
Speed Range
Default 1.0
8+
Voice IDs
Mandarin cast
Capability Highlights

Speech 02 Turbo Speech Architecture

A complete TTS control surface — voice identity, prosody, container format, and delivery mode.

01

Turbo Turnaround

Optimized for interactive voice agents and high-QPS product surfaces that need speech in near real time without dropping persona quality.

Low LatencyInteractiveHigh QPS
Turbo latency
02

Voice Library & Persona

Default voice male-qn-qingse, with male and female Mandarin personas including shaonv, yujie, chengshu, tianmei, jingying, badao, and daxuesheng. Prefer jingying or qingse for crisp assistant tone.

voicemale-qn-qingseMandarin Personas
Voice library
03

Pace & Container Controls

speed maps to voice_setting.speed (default 1.0, range 0.5–2.0). response_format supports mp3, wav, and pcm so post pipelines can pick the right container.

speed 0.5–2.0mp3 / wav / pcmvoice_setting.speed
Pace and format
04

Delivery & Streaming

metadata.output_format switches between url and hex delivery (url default). stream_format enables upstream streaming when any non-empty value is sent.

url / hexstream_formatSSE-ready
Delivery and streaming
How It Works

End-to-End Speech Pipeline

From script to finished, mix-ready audio in five deliberate steps.

01

Write Input Script

Send up to 10,000 characters of punctuated plain text. Written rhythm drives pauses and emphasis.

02

Cast the Voice

Choose a voice_id such as male-qn-qingse, female-yujie, or female-tianmei to match brand persona.

03

Direct Pace

Set speed between 0.5 and 2.0. Slow down for legal or medical copy; speed up for UI microcopy.

04

Pick Container

Select mp3 for direct delivery, wav/pcm for mix and master. Use url or hex via metadata.output_format.

05

Stream or File

Set stream_format when you need progressive chunks; otherwise download the finished asset.

Speech 02 Turbo Production Domains

Where this tier of MiniMax Speech ships every week.

CX

IVR Trees

Menu prompts and hold messages with brand-consistent tone.

IVR
Support

Support Bots

Spoken replies for voice agents and chat escalation.

Voicebot
Ops

Status Broadcasts

Incident and outage reads that stay calm under pressure.

Ops
Product

Appointment Reminders

Short transactional reads with reliable names and times.

Notify
Best Practices

Voice Direction Notes

Studio craft that lifts first-pass quality more than any single parameter.

Punctuate for prosody

Use commas, periods, and line breaks to control pauses — the model follows written rhythm closely.

Match voice to persona

Pick shaonv / tianmei for bright product UI; yujie / chengshu for premium brand VO; daxuesheng for campus tone.

Prefer wav for post

Use wav or pcm when you will mix or master; mp3 for direct CDN delivery to end users.

Read numbers out loud

Write “twenty twenty-four” instead of “2024” when you need spoken years or codes to land cleanly.

Developer Quickstart

Speech 02 Turbo Quickstart

Synthesize speech with explicit voice, pace, and format controls.

curl -X POST "https://api.powertokens.ai/v1/audio/speech" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "speech-02-turbo",
  "input": "欢迎使用 MiniMax 语音合成,这是一段示例播报。",
  "voice": "male-qn-qingse",
  "speed": 1,
  "response_format": "mp3"
}' \
  --output speech.mp3

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
MiniMax • speech-02-turbo
Capability
Text-to-Speech (POST /v1/audio/speech)
Input
input text up to 10,000 characters
Voice
voice default male-qn-qingse (8+ Mandarin personas)
Speed
speed default 1.0, range 0.5–2.0 → voice_setting.speed
Response Format
mp3 (default), wav, pcm
Output Delivery
metadata.output_format: url or hex
Streaming
stream_format non-empty enables upstream stream
Billing
Usage-based via pricing matrix
Family
Speech 02 Turbo

Détails des tarifs

La facturation réelle de ce modèle est calculée dynamiquement en fonction des paramètres spécifiques de votre requête API. Voici les combinaisons spécifiques et leurs tarifs correspondants :

Standard
Crédits57/ 1K Characters
Prix (USD)$0.057
Same Channel

Models from the Same Channel

Explore complementary models and alternative versions from the same provider channel.

MiniMax Speech 2.8 HD
audioCommercial
MiniMax Speech 2.8 HD

MiniMax Speech 2.8 HD

speech-2.8-hd

speech-2.8-hd is a high-definition AI speech synthesis model tailored for individual users. It delivers studio-grade natural voice texture with ultra-realistic pronunciation and smooth intonation. It supports rich exclusive timbres and multilingual conversion, and is capable of simulating vivid emotions like laughter and sighs. It perfectly fits daily voice dubbing, audio creation, reading narration and personal voice customization, bringing you immersive and high-quality voice experience.

Text to Music
MiniMax Speech 2.8 Turbo
audioCommercial
MiniMax Speech 2.8 Turbo

MiniMax Speech 2.8 Turbo

speech-2.8-turbo

speech-2.8-turbo is a lightweight and ultra-fast AI speech synthesis model for all users. It features instant response, efficient generation and stable audio output. With natural and smooth timbre performance, it supports multilingual conversion and basic emotional intonation adjustment. Optimized for low-latency scenarios such as daily narration, short video dubbing and real-time voice interaction, it balances speed, quality and ease of use, delivering a fluent and convenient voice creation exp

Text to Music
MiniMax Speech 2.6 HD
audioCommercial
MiniMax Speech 2.6 HD

MiniMax Speech 2.6 HD

speech-2.6-hd

speech-2.6-hd is a high-definition AI voice synthesis model designed for general users. It delivers lifelike, studio-level vocal quality with natural pronunciation, smooth rhythm and rich emotional expression. It supports multiple languages and diverse premium voice tones, enabling vivid voice dubbing, audiobook narration and personalized voice creation. With stable sound quality and authentic intonation, it perfectly fits daily entertainment, content creation and daily voice playback needs, bri

Text to Music
MiniMax Speech 2.6 Turbo
audioCommercial
MiniMax Speech 2.6 Turbo

MiniMax Speech 2.6 Turbo

speech-2.6-turbo

A lightweight, ultra-fast AI speech synthesis model with premium sound quality and sub-250ms low latency. Supports 40+ languages, 7 emotions, and 300+ curated voices for real-time interaction, short video dubbing, and daily narration. Delivers natural, smooth audio with high cost-performance.

Text to Music

Frequently Asked Questions

Everything you need to know before integrating this model.

Up to 10,000 characters per request in the input field.

Start Building with Speech 02 Turbo Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.