Z.ai's native multimodal agent model for vision-based coding and agent workflows. This is the standard non-thinking variant for image, video, and text inputs, tuned for perceive-plan-execute loops, complex coding, and tool-driven task execution.
Added Apr 1, 2026
Context Window
202.8K
Max Output
131.1K
Avg output tokens (7d)
279 tokens
Input Price (Auto)
$1.20/1M
Output Price (Auto)
$4.00/1M
Cache Read (Auto)
$0.24/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from LMArena.
Arena Score
1433.5
Overall Rank
#104 / 409
Votes
10,393
Confidence Interval
1427.0 - 1439.9
Category Scores
Coding
#93 / 404
2,923 votes
1488.1
Math
#88 / 393
480 votes
1436.8
Longer Query
#104 / 387
4,878 votes
1441.8
Creative Writing
#98 / 407
1,901 votes
1403.0
Instruction Following
#103 / 409
3,690 votes
1423.0
Hard Prompts
#106 / 409
6,935 votes
1452.2
Additional Categories21
English
#82 / 409
4,351 votes
1452.0
Spanish
#82 / 292
307 votes
1437.7
Korean
#84 / 276
220 votes
1388.1
Hard Prompts English
#89 / 407
2,918 votes
1466.4
Industry Mathematical
#90 / 390
599 votes
1445.1
Expert
#95 / 359
1,200 votes
1461.7
Polish
#96 / 229
197 votes
1425.6
Industry Writing And Literature And Language
#97 / 408
2,706 votes
1412.2
Chinese
#99 / 385
774 votes
1473.9
Industry Software And It Services
#99 / 409
4,141 votes
1472.7
Industry Business And Management And Financial Operations
#100 / 402
2,141 votes
1439.7
Exclude Ties
#105 / 409
7,662 votes
1427.0
French
#105 / 287
419 votes
1444.4
Russian
#105 / 373
1,118 votes
1427.8
Multi Turn
#107 / 407
1,729 votes
1433.5
Non English
#107 / 409
6,040 votes
1414.0
Industry Entertainment And Sports And Media
#110 / 407
2,488 votes
1396.3
Industry Life And Physical And Social Science
#110 / 407
1,651 votes
1447.6
German
#112 / 306
178 votes
1411.0
Industry Legal And Government
#119 / 382
825 votes
1431.1
Industry Medicine And Healthcare
#132 / 379
744 votes
1433.5
Published 2026-09-25 · Matched as glm-5v-turbo
LMArena DatasetProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM 5V Turbo with similar models from the same provider or model family.
GLM 5V Turbo Thinking
z-ai/glm-5v-turbo:thinkingThinking-enabled GLM 5V Turbo for image, video, and text inputs. Uses the same multimodal foundation model with more deliberate vision-grounded analysis, planning, and tool use.
GLM 5 Turbo
z-ai/glm-5-turboFast GLM 5 Turbo variant from Z-AI for general chat, coding, and tool use.
GLM 4.6 Turbo
z-ai/GLM-4.6-turboFast variant of GLM 4.6 for general chat, coding, and analysis with improved latency and strong reasoning.
GLM 4.6 Turbo (Thinking)
z-ai/GLM-4.6-turbo:thinkingGLM 4.6 Turbo with thinking mode enabled for enhanced reasoning; shows internal reasoning and supports long context.
GLM 5.3 Flash Cybersecurity
z-ai/glm-5.3-flash-cybersecurityGLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.