Browse all Z.AI text models
Provider logo

GLM 5 NVFP4 TEE

z-ai/GLM-5-NVFP4-TEE
Back
Provider logo

GLM 5 NVFP4 TEE

z-ai/GLM-5-NVFP4-TEE
Back

GLM 5 (NVFP4-quantized) running through Chutes TEE with encrypted transport and provider evidence support.

Added Jun 10, 2026

Context Window

202.8K

Max Output

65.5K

Input Price (Auto)

$0.49/1M

Output Price (Auto)

$1.96/1M

Cache Read (Auto)

$0.24/1M

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

21.8

Better than 69% of models compared

Agentic work

T²-Bench Telecom (legacy)

Legacy fallback · Conversational AI agents in dual-control scenarios

97.4%

Better than 97% of models compared

Document reasoning

AA-LCR v1.1

Long context reasoning with updated grading

43.7%

Better than 39% of models compared

Reasoning

HLE

Humanity's Last Exam

7.6%

Better than 44% of models compared

IFBench

Instruction-following benchmark

55.2%

Better than 66% of models compared

Coding

Terminal-Bench Hard (legacy)

Legacy fallback · Agentic coding and terminal use

39.4%

Better than 86% of models compared

Legacy benchmarks

GPQA Diamond (legacy)

Graduate-level scientific reasoning

66.6%

Better than 43% of models compared

AA-LCR (unversioned / legacy)

Long context reasoning evaluation

43.7%

Better than 39% of models compared

Last updated Oct 1, 2026

Artificial Analysis

Providers

Provider information for this model’s automatic routing. These routes cannot be selected individually.

Loading provider options…

Compare GLM 5 NVFP4 TEE with similar models from the same provider or model family.

GLM 5.3 Flash Cybersecurity

z-ai/glm-5.3-flash-cybersecurity

GLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.

GLM 5.3 Flash

z-ai/glm-5.3-flash

ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM 5.3

z-ai/glm-5.3

GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.

GLM 5.3 Thinking

z-ai/glm-5.3:thinking

GLM-5.3 with higher reasoning enabled for harder long-horizon coding, autonomous agent workflows, and complex engineering tasks.

GLM 5.3 Flash Uncensored

z-ai/glm-5.3-flash-uncensored

GLM 5.3 Flash Uncensored is an uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with provider-dependent vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.

GLM 5.2

z-ai/glm-5.2

GLM-5.2 is Z.AI's flagship model for long-horizon autonomous coding and engineering workflows. It is built to plan, execute, iterate, and optimize complex development tasks over extended runs. This variant keeps thinking disabled for faster direct responses.