Google's Gemma 4 31B instruction-tuned model for heavier reasoning, coding, agentic workflows, and long-context multimodal understanding. This route keeps tokenizer thinking disabled for faster direct answers. Requests containing video cost $0.14 per million input tokens and $0.40 per million output tokens.
Added Apr 2, 2026
Model weightsContext Window
262.1K
Max Output
131.1K
Avg output tokens (7d)
291 tokens
Input Price (Auto)
$0.14/1M
Output Price (Auto)
$0.40/1M
Cache Read (Auto)
$0.070/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
14.7
Coding Index
43.4
Agentic Index
4.2
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
4.9%
Better than 28% of models compared
Harvey LAB-AA
Legal agentic work criterion pass rate
47.2%
Better than 4% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
365 Elo
Better than 19% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
614 Elo
Better than 31% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
6.0%
Better than 30% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
69.7%
Better than 62% of models compared
Reasoning
HLE
Humanity's Last Exam
23.6%
Better than 72% of models compared
IFBench
Instruction-following benchmark
75.6%
Better than 93% of models compared
CritPt
Research-level physics reasoning
1.4%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 16% of models compared
SciCode
Python programming for scientific computing
45.5%
Better than 41% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
20.0%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.0%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
85.7%
Better than 79% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
36.4%
Better than 83% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
59.9%
Better than 56% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
69.7%
Better than 62% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
5.7%
Last updated Oct 2, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Gemma 4 31B with similar models from the same provider or model family.
Gemma 4 31B Thinking
google/gemma-4-31b-it:thinkingGoogle's Gemma 4 31B instruction-tuned model with thinking explicitly enabled, exposing reasoning traces for complex multimodal and coding workflows. Requests containing video cost $0.14 per million input tokens and $0.40 per million output tokens.
Gemma 4 26B A4B Cybersecurity
google/gemma-4-26b-a4b-it-cybersecurityGemma 4 26B A4B Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports optional reasoning, image understanding, tool calling, and a 262,144-token context window.
Gemma 4 26B A4B
google/gemma-4-26b-a4b-itGoogle's Gemma 4 26B A4B instruction-tuned model built for scalable reasoning, coding, long-context, and multimodal workflows. This route is tuned for faster direct answers while preserving multimodal and structured output support. Requests containing video cost $0.10 per million input tokens and $0.40 per million output tokens.
Gemma 4 26B A4B Thinking
google/gemma-4-26b-a4b-it:thinkingGoogle's Gemma 4 26B A4B instruction-tuned model with structured reasoning for more deliberate coding, multimodal analysis, and long-context problem solving. Requests containing video cost $0.10 per million input tokens and $0.40 per million output tokens.
DiffusionGemma
google/diffusiongemmaDiffusionGemma is a high-speed diffusion-based version of Gemma 4 26B A4B. It supports optional reasoning and a 262,144-token context window.
Gemma 4 31B Split-Untied
google/gemma4-31b-splituntiedBlazed-Forge's Split-Untied is a text-only Gemma 4 31B community finetune with an untied BF16 output head, built for creative writing, roleplay, expressive dialogue, and tool use.