
DeepSeek V4 Pro / Chat
Large-scale MoE model with 1.6T total parameters and 49B activated, supporting a 1M-token context window . Designed for advanced reasoning, coding, and long-horizon agent workflows . Top performance on GPQA Diamond (88.8%) and Terminal-Bench Hard (46.2%)
Press Enter to add, Backspace to remove
Playground Chat
Starten Sie eine Unterhaltung mit dem KI-Modell. Sie können alles fragen.
KI-generierte Antworten können in der Genauigkeit variieren.
Preisdetails
Die tatsächliche Abrechnung für dieses Modell wird dynamisch basierend auf den spezifischen Parametern Ihrer API-Anfrage berechnet. Nachfolgend finden Sie die spezifischen Kombinationen und ihre entsprechenden Preise:
- Für Anfragen an Werktagen (Mo–Fr) zwischen 09:00 - 12:00, 14:00 - 18:00 (UTC+8) wird ein Preismultiplikator von 2x angewendet.
| Modalität | Eingabe-Guthaben | Ausgabe-Guthaben | Eingabepreis | Ausgabepreis | Impliziter Cache-Treffer | Expliziter Cache-Treffer | Cache-Erstellung |
|---|---|---|---|---|---|---|---|
| Standard | 660/ 1M Tokens | 1,980/ 1M Tokens | $0.660 | $1.980 | $0.022 22/ 1M Tokens | $0.022 22/ 1M Tokens | -- |
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


GLM-5.3
glm-5.3
Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)


Qwen3.8 Flash
qwen3.8-flash
Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration


GLM-5.2
glm-5.2
Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips


Qwen3 Max
qwen3-max
Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely