
MiniMax Speech 2.8 HD / テキストから音声
speech-2.8-hd is a high-definition AI speech synthesis model tailored for individual users. It delivers studio-grade natural voice texture with ultra-realistic pronunciation and smooth intonation. It supports rich exclusive timbres and multilingual conversion, and is capable of simulating vivid emotions like laughter and sighs. It perfectly fits daily voice dubbing, audio creation, reading narration and personal voice customization, bringing you immersive and high-quality voice experience.
料金詳細
このモデルの実際の課金は、API リクエストで渡される特定のパラメータに基づいて動的に計算されます。以下は具体的な組み合わせとそれに対応する料金です:
| モダリティ | クレジット | 料金 (USD) |
|---|---|---|
| Standard | 95/ 1K Characters | $0.095 |
Models from the Same Channel
Explore complementary models and alternative versions from the same provider channel.


MiniMax Speech 2.8 Turbo
speech-2.8-turbo
speech-2.8-turbo is a lightweight and ultra-fast AI speech synthesis model for all users. It features instant response, efficient generation and stable audio output. With natural and smooth timbre performance, it supports multilingual conversion and basic emotional intonation adjustment. Optimized for low-latency scenarios such as daily narration, short video dubbing and real-time voice interaction, it balances speed, quality and ease of use, delivering a fluent and convenient voice creation exp


MiniMax Speech 2.6 HD
speech-2.6-hd
speech-2.6-hd is a high-definition AI voice synthesis model designed for general users. It delivers lifelike, studio-level vocal quality with natural pronunciation, smooth rhythm and rich emotional expression. It supports multiple languages and diverse premium voice tones, enabling vivid voice dubbing, audiobook narration and personalized voice creation. With stable sound quality and authentic intonation, it perfectly fits daily entertainment, content creation and daily voice playback needs, bri


MiniMax Speech 2.6 Turbo
speech-2.6-turbo
A lightweight, ultra-fast AI speech synthesis model with premium sound quality and sub-250ms low latency. Supports 40+ languages, 7 emotions, and 300+ curated voices for real-time interaction, short video dubbing, and daily narration. Delivers natural, smooth audio with high cost-performance.


MiniMax Speech-02 HD
speech-02-hd
A high-definition AI speech model with outstanding prosody, stability, and industry-leading voice cloning similarity. Generates studio-grade, ultra-natural voices with rich emotional expression. Ideal for audiobooks, premium voiceovers, and professional content creation in 40+ languages.
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


Qwen3 TTS Instruct Flash
qwen3-tts-instruct-flash
Qwen3-TTS-Flash model is Tongyi's latest real-time speech synthesis model. The Instruct model processes the synthesis effect through natural language, ensuring highly appropriate emotional and expressive speech in different contexts. Currently, it supports 25 timbres for both Chinese and English Instruct adjustments.


Vidu Audio 1.0
audio1.0
Vidu's text-to-audio generation model (model ID: audio1.0). Generates sound effects and background music from text prompts. Duration range: 2–10 seconds. Supports random seed configuration for reproducible outputs