Get 100 free credits on sign up to explore and build your AI applicationsClaim Free
glm-5

GLM-5 / Chat

Commercial
ID: glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading closed-source models. With advanced agentic planning, deep backend reasoning, and iterative self-correction, GLM-5 moves beyond code generation to full-system construction and autonomous execution.

Input$0.97/ 1M Tokens(970 Credits)
Output$3.104/ 1M Tokens(3104 Credits)
1
0.95

Press Enter to add, Backspace to remove

Chat Conversation

Playground Chat

Start a conversation with the AI model. You can ask anything.

0

AI-generated responses may vary in accuracy.

Zhipu GLM-5 • Thinking Chat

GLM-5Toggleable Deep Thinking with Tools

GLM-5 is Zhipu’s thinking-capable chat model with thinking.type enable/disable, clear_thinking history cleanup, up to 128 function tools, and streaming chat completions on an OpenAI-compatible surface.

thinking.type Toggle
clear_thinking Cleanup
Function Tools (≤128)
Streaming Default On
max_tokens ≤ 131072
ByteDance Seed Foundation Architecture
Commercial License & Enterprise SLA
GLM-5
Zhipu Tier
Toggle
Deep Thinking
GLM-5 Hero

At a Glance

GLM-5
Zhipu Tier
Thinking chat
Toggle
Deep Thinking
thinking.type
128
Max Tools
function only
Stream
Default On
SSE tokens
Capability Highlights

GLM-5 Thinking Stack

Deep reasoning when you need it — plain chat when you do not. One model, two cognitive gears.

Thinking toggle

Thinking Toggle

thinking.type switches between enabled and disabled so you can trade analytical depth for latency per request without swapping models.

thinking.typeenabled / disabledPer Request
Clear thinking

Clear Thinking History

thinking.clear_thinking (default true) cleans historical reasoning_content so multi-turn context stays compact and free of stale intermediate traces.

clear_thinkingDefault TrueMulti-turn
Function Tools

Function Tools

Register up to 128 function tools with tool_choice auto for agent loops, structured actions, and grounded retrieval handoffs.

tools ≤ 128tool_choice: autoFunctions
Sampling & Stop

Sampling & Stop

Temperature [0,1] default 1, top_p [0.01,1] default 0.95, max_tokens up to 131072, and a single stop word for clean parser boundaries.

temperaturetop_pstop (1 word)
How It Works

How It Works

From request to reasoned answer in four moves.

01

Set Thinking

Enable or disable thinking.type based on task complexity and latency budget.

02

Attach Tools

Register function tools (up to 128) with tool_choice auto for agent loops.

03

Stream Tokens

stream defaults to true and returns text/event-stream for progressive paint.

04

Shape Output

Tune temperature (0–1, default 1), top_p (0.01–1, default 0.95), and one stop word.

GLM-5 Domains

Where thinking depth is a product dial, not a model swap.

Analytics

Complex Analysis

Thinking on for multi-hop technical reasoning.

Thinking On
Agents

Agent Workflows

Function tools with tool_choice auto.

Tools
CX

Fast Chat UX

Thinking off for lightweight conversational turns.

Thinking Off
Platform

Structured Output

Temperature 0–1 and single stop word for parsers.

Stop
Best Practices

Prompt Tips

Get better first-pass results from GLM-5.

Turn thinking on for hard tasks

Leave thinking enabled for proofs and multi-document work; disable it for simple labels and rewrites.

Keep stop to one token

GLM-5 currently accepts a single stop word — design parsers around that limit.

Request structured artifacts

Ask for JSON or numbered steps when you will parse output programmatically.

Developer Quickstart

GLM-5 Quickstart

Thinking chat with function tools.

curl -X POST "https://api.powertokens.ai/v1/chat/completions" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5",
    "messages": [
      {"role": "user", "content": "Outline a careful multi-step plan to migrate a REST API to versioned endpoints."}
    ],
    "stream": true,
    "temperature": 0.7
  }'

Technical Specifications

Confirmed parameters and runtime execution protocols.

Provider & Model ID
Zhipu • glm-5
Thinking
thinking.type enabled / disabled (default enabled)
Clear Thinking
thinking.clear_thinking default true
Streaming
stream default true (text/event-stream)
Sampling
Temperature [0,1] default 1; top_p [0.01,1] default 0.95
Output Limits
max_tokens 1–131072
Stop
stop — single stop word only (maxItems 1)
Tools
tools up to 128 functions; tool_choice auto
Protocol
OpenAI-compatible chat completions
API Endpoint
POST /v1/chat/completions

Pricing Details

The actual billing for this model is dynamically calculated based on the specific parameters passed in your API request. Below are the specific combinations and their corresponding pricing:

Standard
Input Price
$0.970(970 / 1M Tokens)
Output Price
$3.104(3,104 / 1M Tokens)
Implicit Cache Hit
$0.194(194/ 1M Tokens)
Explicit Cache Hit
$0.194(194/ 1M Tokens)
Cache Creation
--
Same Channel

Models from the Same Channel

Explore complementary models and alternative versions from the same provider channel.

GLM-5.3
chatCommercial
GLM-5.3

GLM-5.3

glm-5.3

Zhipu AI's flagship text model optimized for complex software engineering and long-horizon Agent tasks. Features a 1M-token context window with mandatory reasoning (3 levels: low/high/max). Coding capability improved 50% over GLM-5.2 on Z.ai Code Bench; scores SOTA on Terminal Bench 3.0. Excels in cybersecurity tasks (cyber vulnerability discovery)

ChatText Generation
GLM-5.2
chatCommercial
GLM-5.2

GLM-5.2

glm-5.2

Flagship text model purpose-built for long-horizon agentic workflows. Features a 1M context window supporting project-level engineering in a single session. Excels at autonomous coding: can complete development, testing, and multi-platform deployment from a single prompt. Top open-weight model per Artificial Analysis; #1 globally on Code Arena. MIT-licensed and Day-0 optimized for domestic AI chips

ChatText Generation
GLM-5 Turbo
chatCommercial
GLM-5 Turbo

GLM-5 Turbo

glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, with improved complex instruction decomposition, tool use, scheduled and persistent execution, and overall stability across extended tasks.

Chat
GLM-5.1
chatCommercial
GLM-5.1

GLM-5.1

glm-5.1

High-performance text model delivering breakthrough coding and long-horizon task execution. Capable of autonomous, continuous work for 8+ hours per session—planning, executing, and iterating to deliver engineering-grade results. Coding capability aligns with Claude Opus 4.6; scores 58.4 on SWE-Bench Pro, surpassing GPT-5.4 and Opus 4.6. 200K context window. Optimized for Agentic Coding, MCP tool calling, and complex software engineering

ChatText Generation
Ecosystem Models

Recommended Related Models

Explore complementary video and multimodal models with your unified API key.

Browse All Models
Qwen3.8 Flash
chatCommercial
Qwen3.8 Flash

Qwen3.8 Flash

qwen3.8-flash

Alibaba's cost-efficient multimodal reasoning model. Supports text, image, and video inputs with text output. Features a native 1M-token context window for long documents, codebases, and agentic workflows. Excels in coding assistance, desktop interaction, chart analysis, and long-video understanding. Compatible with OpenAI/Anthropic protocols for seamless integration

ChatText Generation
Qwen3 Max
chatCommercial
Qwen3 Max

Qwen3 Max

qwen3-max

Compared with the September 23, 2025 version, the newly upgraded Qwen-3 Max seamlessly integrates thinking and non-thinking modes, bringing an all-round obvious performance boost. Its thinking mode supports web search, web content extraction and code interpreter. It can conduct in-depth logical reasoning and call external tools to solve intricate problems more precisely

Chat
MiniMax M3
chatCommercial
MiniMax M3

MiniMax M3

MiniMax-M3

Flagship multimodal foundation model supporting text, image, and video inputs with text output. Features a 1M-token context window via MiniMax Sparse Attention (MSA), cutting per-token compute to ~1/20 of previous gen at full context. Excels at long-horizon agentic work, coding, and tool use. Native multimodal training from step zero ensures deep semantic alignment. Scores 59.0% on SWE-Bench Pro and 83.5 on BrowseComp, surpassing Opus 4.7

ChatText Generation
Qwen3 Coder Plus
chatCommercial
Qwen3 Coder Plus

Qwen3 Coder Plus

qwen3-coder-plus

Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.

Chat

Frequently Asked Questions

Everything you need to know before integrating this model.

Set thinking.type to enabled or disabled. Default is enabled.

Start Building with GLM-5 Today

Create an account in seconds to receive 100 free credits and start generating immediately. No credit card or upfront contract required.