
DeepSeek V4 Pro / Chat
Large-scale MoE model with 1.6T total parameters and 49B activated, supporting a 1M-token context window . Designed for advanced reasoning, coding, and long-horizon agent workflows . Top performance on GPQA Diamond (88.8%) and Terminal-Bench Hard (46.2%)
Press Enter to add, Backspace to remove
Playground Chat
Start a conversation with the AI model. You can ask anything.
AI-generated responses may vary in accuracy.
Pricing Details
The actual billing for this model is dynamically calculated based on the specific parameters passed in your API request. Below are the specific combinations and their corresponding pricing:
- For requests made on weekdays (Mon-Fri) during 09:00 - 12:00, 14:00 - 18:00 (UTC+8), a pricing multiplier of 2x is applied.
| Modality | Input Credits | Output Credits | Input Price | Output Price | Implicit Cache Hit | Explicit Cache Hit | Cache Creation |
|---|---|---|---|---|---|---|---|
| Standard | 660/ 1M Tokens | 1,980/ 1M Tokens | $0.660 | $1.980 | $0.022 22/ 1M Tokens | $0.022 22/ 1M Tokens | -- |
Recommended Related Models
Explore complementary video and multimodal models with your unified API key.


GLM-5.3
glm-5.3
Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)


Qwen3.8 Flash
qwen3.8-flash
Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration


GLM-5.2
glm-5.2
Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips


Qwen3 Max
qwen3-max
Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely