GLM 5 (NVFP4-quantized) running through Chutes TEE with encrypted transport and provider evidence support.
Added Jun 10, 2026
Context Window
202.8K
Max Output
65.5K
Input Price (Auto)
$0.49/1M
Output Price (Auto)
$1.96/1M
Cache Read (Auto)
$0.24/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
21.8
Agentic work
T²-Bench Telecom (legacy)
Legacy fallback · Conversational AI agents in dual-control scenarios
97.4%
Better than 97% of models compared
Document reasoning
AA-LCR v1.1
Long context reasoning with updated grading
43.7%
Better than 39% of models compared
Reasoning
HLE
Humanity's Last Exam
7.6%
Better than 44% of models compared
IFBench
Instruction-following benchmark
55.2%
Better than 66% of models compared
Coding
Terminal-Bench Hard (legacy)
Legacy fallback · Agentic coding and terminal use
39.4%
Better than 86% of models compared
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
66.6%
Better than 43% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
43.7%
Better than 39% of models compared
Last updated Oct 1, 2026
Artificial AnalysisProviders
Provider information for this model’s automatic routing. These routes cannot be selected individually.
Loading provider options…
Related text models
Compare GLM 5 NVFP4 TEE with similar models from the same provider or model family.
GLM 5.3 Flash Cybersecurity
z-ai/glm-5.3-flash-cybersecurityGLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM 5.3
z-ai/glm-5.3GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.
GLM 5.3 Thinking
z-ai/glm-5.3:thinkingGLM-5.3 with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks.
GLM 5.3 Flash Uncensored
z-ai/glm-5.3-flash-uncensoredGLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with provider-dependent vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
GLM 5.2
z-ai/glm-5.2GLM-5.2 is Z.AI's flagship model for long-horizon autonomous coding and engineering workflows. It is built to plan, execute, iterate, and optimize complex development tasks over extended runs. This variant keeps thinking disabled for faster direct responses.