
MiniMax M2.7 Highspeed / Chat
A high-speed variant of MiniMax’s flagship M2.7 LLM, delivering 100 TPS—~60% faster than standard M2.7—with identical top-tier quality. It features a 204,800-token context window, near-Opus-level SWE performance, and recursive self-improvement. Optimized for low-latency coding, office tasks, and agent workflows, it balances extreme speed, reliability, and cost-efficiency for high-throughput real-world applications。
Press Enter to add, Backspace to remove
Playground Chat
Inicie uma conversa com o modelo de IA. Você pode perguntar qualquer coisa.
As respostas geradas por IA podem variar em precisão.

MiniMax M2.7 High SpeedProduction Chat at Interactive Latency
MiniMax M2.7 High Speed is a latency-optimized chat sibling of M2.7 with streaming defaults, temperature and top-p sampling, stop sequences, and max output tokens up to 2048 — for production chat that must stay interactive under load.
At a Glance
M2.7 High Speed Capabilities
Same chat completions contract as M2.7, tuned for interactive latency at production QPS.
High-Speed Serving Path
A dedicated low-latency path for production chat surfaces that need immediate first tokens under load — support consoles, search answers, and agent loops.

Streaming by Default
stream defaults to true so tokens arrive as they are generated. Ideal for long answers where partial paint keeps users engaged.

Sampling Controls
Temperature default 0.7 (range 0–1) and top-p default 0.95 (range 0–1) for balanced creative range with stable brand tone.

Bounded Completions
Stop sequences and max_tokens (1–2048) keep completions inside product budgets and schema boundaries even when prompts drift.

How It Works
Simple high-speed chat completions path.
POST Messages
Send system and user turns in the standard chat completions schema.
Stream Tokens
Consume SSE for immediate UI paint across long answers.
Bound Output
Use stop sequences and max_tokens (1–2048) for parser-safe replies.
Tune Sampling
Adjust temperature and top-p for brand voice across tenants.
High Speed Product Domains
Where production chat must stay interactive.
Customer Chat
Streaming support replies with stop-bounded answers.
Content Assist
Fast drafts, rewrites, and summaries.
Ops Automation
Ticket drafting and internal Q&A at scale.
Search Answers
Grounded replies over retrieved context.
Prompt Tips
Keep M2.7 High Speed responses reliable.
Put role, constraints, and output format in the system message.
Keep stream on for chat surfaces; disable only for batch jobs.
Bound generation when you will parse JSON or field lists.
MiniMax M2.7 High Speed Quickstart
Low-latency streaming production chat.
curl -X POST "https://api.powertokens.ai/v1/chat/completions" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMax-M2.7-highspeed",
"messages": [
{"role": "user", "content": "Draft a concise out-of-office reply for a client email."}
],
"stream": true,
"temperature": 0.7
}'Technical Specifications
Confirmed parameters and runtime execution protocols.
Detalhes de preços
A cobrança real deste modelo é calculada dinamicamente com base nos parâmetros específicos da sua solicitação de API. Abaixo estão as combinações específicas e seus preços correspondentes:
| Modalidade | Créditos de entrada | Créditos de saída | Preço de entrada | Preço de saída | Cache implícito | Cache explícito | Criação de cache |
|---|---|---|---|---|---|---|---|
| Standard | 600/ 1M Tokens | 2,400/ 1M Tokens | $0.600 | $2.400 | $0.060 60/ 1M Tokens | $0.060 60/ 1M Tokens | $0.375 375/ 1M Tokens |
Models from the Same Channel
Explore complementary models and alternative versions from the same provider channel.


MiniMax M3
MiniMax-M3
Flagship multimodal foundation model supporting text, image, and video inputs with text output. Features a 1M-token context window via MiniMax Sparse Attention (MSA), cutting per-token compute to ~1/20 of previous gen at full context. Excels at long-horizon agentic work, coding, and tool use. Native multimodal training from step zero ensures deep semantic alignment. Scores 59.0% on SWE-Bench Pro and 83.5 on BrowseComp, surpassing Opus 4.7


MiniMax M2.7
MiniMax-M2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent collaboration, enabling it to plan, execute, and refine complex tasks across dynamic environments.


MiniMax M2.5
MiniMax-M2.5
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams


MiniMax M2.5 Highspeed
MiniMax-M2.5-highspeed
MiniMax-M2.5-highspeed is an ultra-fast inference version of MiniMax M2.5. It maintains the full intelligent capability of the standard M2.5 model, featuring ultra-low latency, high throughput and outstanding cost performance. Built on advanced MoE architecture, it delivers rapid response speed while ensuring stable output quality. Optimized for high-concurrency business scenarios such as real-time dialogue, content generation and API service calls, it perfectly meets enterprise-level demands fo
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


GLM-5.3
glm-5.3
Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)


Qwen3.8 Flash
qwen3.8-flash
Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration


GLM-5.2
glm-5.2
Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips


Qwen3 Max
qwen3-max
Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely
Frequently Asked Questions
Everything you need to know before integrating this model.
Start Building with MiniMax M2.7 High Speed Today
Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.