新用户免费领取 100 积分,即刻探索与构建您的 AI 应用免费领取
speech-02-hd

MiniMax Speech-02 HD / 文本转语音

Commercial
ID: speech-02-hd

A high-definition AI speech model with outstanding prosody, stability, and industry-leading voice cloning similarity. Generates studio-grade, ultra-natural voices with rich emotional expression. Ideal for audiobooks, premium voiceovers, and professional content creation in 40+ languages.

价格$0.095/ 1K Characters(95 积分)
327 total
1
语音合成 (TTS)
0 / 10,000
音频输出准备就绪。点击「开始合成」生成语音。
0:00 / 0:00
MiniMax Speech 02 HD • Text-to-Speech

Speech 02 HDHigh-fidelity studio TTS

Speech 02 HD is MiniMax 02-series speech synthesis with selectable Mandarin voices, speed control, and mp3 / wav / pcm output — tuned for high-fidelity studio narration. Prefer yujie or chengshu for premium brand VO.

HD Voice Fidelity
8 Core Voice IDs
Speed 0.5–2.0
mp3 / wav / pcm
url / hex Delivery
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
Speech 02 HD Hero

At a Glance

02
Series
MiniMax Speech 02
HD
Tier
High fidelity
10K
Input Chars
Per request
3
Formats
mp3 · wav · pcm
0.5–2.0
Speed Range
Default 1.0
8+
Voice IDs
Mandarin cast
Capability Highlights

Speech 02 HD Speech Architecture

A complete TTS control surface — voice identity, prosody, container format, and delivery mode.

HD Voice Fidelity

Prioritizes natural prosody, breath, and studio-grade clarity for narration, IVR, and brand voice assets that ship as finished media.

HD FidelityStudioNarration
HD voice fidelity
Voice library

Voice Library & Persona

Default voice male-qn-qingse, with male and female Mandarin personas including shaonv, yujie, chengshu, tianmei, jingying, badao, and daxuesheng. Prefer yujie or chengshu for premium brand VO.

voicemale-qn-qingseMandarin Personas

Pace & Container Controls

speed maps to voice_setting.speed (default 1.0, range 0.5–2.0). response_format supports mp3, wav, and pcm so post pipelines can pick the right container.

speed 0.5–2.0mp3 / wav / pcmvoice_setting.speed
Pace and format
Delivery and streaming

Delivery & Streaming

metadata.output_format switches between url and hex delivery (url default). stream_format enables upstream streaming when any non-empty value is sent.

url / hexstream_formatSSE-ready
How It Works

End-to-End Speech Pipeline

From script to finished, mix-ready audio in five deliberate steps.

01

Write Input Script

Send up to 10,000 characters of punctuated plain text. Written rhythm drives pauses and emphasis.

02

Cast the Voice

Choose a voice_id such as male-qn-qingse, female-yujie, or female-tianmei to match brand persona.

03

Direct Pace

Set speed between 0.5 and 2.0. Slow down for legal or medical copy; speed up for UI microcopy.

04

Pick Container

Select mp3 for direct delivery, wav/pcm for mix and master. Use url or hex via metadata.output_format.

05

Stream or File

Set stream_format when you need progressive chunks; otherwise download the finished asset.

Speech 02 HD Production Domains

Where this tier of MiniMax Speech ships every week.

Media

Brand Narration

Premium VO for product films, explainers, and launch films.

VO
Publishing

Audiobook Drafts

Long-form input up to 10,000 characters per synthesis call.

10K chars
Audio

Podcast Intros

Consistent host identity across weekly episode openers.

Series
Education

Course Modules

Clear instructional reads with stable pacing and diction.

Edu
Best Practices

Voice Direction Notes

Studio craft that lifts first-pass quality more than any single parameter.

Punctuate for prosody

Use commas, periods, and line breaks to control pauses — the model follows written rhythm closely.

Match voice to persona

Pick shaonv / tianmei for bright product UI; yujie / chengshu for premium brand VO; daxuesheng for campus tone.

Prefer wav for post

Use wav or pcm when you will mix or master; mp3 for direct CDN delivery to end users.

Read numbers out loud

Write “twenty twenty-four” instead of “2024” when you need spoken years or codes to land cleanly.

Developer Quickstart

Speech 02 HD Quickstart

Synthesize speech with explicit voice, pace, and format controls.

curl -X POST "https://api.powertokens.ai/v1/audio/speech" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "speech-02-hd",
  "input": "欢迎使用 MiniMax 语音合成,这是一段示例播报。",
  "voice": "male-qn-qingse",
  "speed": 1,
  "response_format": "mp3"
}' \
  --output speech.mp3

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
MiniMax • speech-02-hd
Capability
Text-to-Speech (POST /v1/audio/speech)
Input
input text up to 10,000 characters
Voice
voice default male-qn-qingse (8+ Mandarin personas)
Speed
speed default 1.0, range 0.5–2.0 → voice_setting.speed
Response Format
mp3 (default), wav, pcm
Output Delivery
metadata.output_format: url or hex
Streaming
stream_format non-empty enables upstream stream
Billing
Usage-based via pricing matrix
Family
Speech 02 HD

价格详情

此模型的实际计费根据您在 API 请求中传入的具体参数动态计算。以下是具体的组合及其对应的价格:

Standard
积分95/ 1K Characters
价格 (USD)$0.095
Same Channel

Models from the Same Channel

Explore complementary models and alternative versions from the same provider channel.

MiniMax Speech 2.8 HD
audioCommercial
MiniMax Speech 2.8 HD

MiniMax Speech 2.8 HD

speech-2.8-hd

speech-2.8-hd is a high-definition AI speech synthesis model tailored for individual users. It delivers studio-grade natural voice texture with ultra-realistic pronunciation and smooth intonation. It supports rich exclusive timbres and multilingual conversion, and is capable of simulating vivid emotions like laughter and sighs. It perfectly fits daily voice dubbing, audio creation, reading narration and personal voice customization, bringing you immersive and high-quality voice experience.

Text to Music
MiniMax Speech 2.8 Turbo
audioCommercial
MiniMax Speech 2.8 Turbo

MiniMax Speech 2.8 Turbo

speech-2.8-turbo

speech-2.8-turbo is a lightweight and ultra-fast AI speech synthesis model for all users. It features instant response, efficient generation and stable audio output. With natural and smooth timbre performance, it supports multilingual conversion and basic emotional intonation adjustment. Optimized for low-latency scenarios such as daily narration, short video dubbing and real-time voice interaction, it balances speed, quality and ease of use, delivering a fluent and convenient voice creation exp

Text to Music
MiniMax Speech 2.6 HD
audioCommercial
MiniMax Speech 2.6 HD

MiniMax Speech 2.6 HD

speech-2.6-hd

speech-2.6-hd is a high-definition AI voice synthesis model designed for general users. It delivers lifelike, studio-level vocal quality with natural pronunciation, smooth rhythm and rich emotional expression. It supports multiple languages and diverse premium voice tones, enabling vivid voice dubbing, audiobook narration and personalized voice creation. With stable sound quality and authentic intonation, it perfectly fits daily entertainment, content creation and daily voice playback needs, bri

Text to Music
MiniMax Speech 2.6 Turbo
audioCommercial
MiniMax Speech 2.6 Turbo

MiniMax Speech 2.6 Turbo

speech-2.6-turbo

A lightweight, ultra-fast AI speech synthesis model with premium sound quality and sub-250ms low latency. Supports 40+ languages, 7 emotions, and 300+ curated voices for real-time interaction, short video dubbing, and daily narration. Delivers natural, smooth audio with high cost-performance.

Text to Music

Frequently Asked Questions

Everything you need to know before integrating this model.

Up to 10,000 characters per request in the input field.

Start Building with Speech 02 HD Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.