The thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.
Added Jul 15, 2026
Model weightsContext Window
1.0M
Max Output
32.8K
Avg output tokens (7d)
366 tokens
Input Price (Auto)
$1.00/1M
Output Price (Auto)
$4.05/1M
Cache Read (Auto)
$0.17/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
25.0
Coding Index
52.1
Agentic Index
22.5
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
5.0%
Better than 29% of models compared
AutomationBench-AA Tasks Completed
Fully completed workflows without guardrail violations
0.3%
Better than 3% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
834 Elo
Better than 36% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
1064 Elo
Better than 53% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
12.8%
Better than 51% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
77.3%
Better than 78% of models compared
MLCR-AA
Medical long-context reasoning
12.2%
Better than 38% of models compared
Reasoning
HLE
Humanity's Last Exam
31.9%
Better than 80% of models compared
CritPt
Research-level physics reasoning
5.4%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
1.0%
Better than 42% of models compared
SciCode
Python programming for scientific computing
47.0%
Better than 44% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
41.5%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
67.7%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
87.2%
Better than 83% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
77.3%
Better than 78% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
28.2%
Last updated Oct 1, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Inkling Thinking with similar models from the same provider or model family.
Inkling
thinkingmachines/inklingThe non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.
Inkling Small
thinkingmachines/Inkling-SmallThe direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
thinkingmachines/Inkling-Small:thinkingThe reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Claude Opus 5.5
anthropic/claude-opus-5.5Claude Opus 5.5 is Anthropic's flagship model for agentic coding, computer use, complex knowledge work, and long-running tasks, with improved efficiency and more natural communication.
Nex N2.5 Mini
nex-agi/nex-n2.5-miniNex N2.5 Mini is an open-source multimodal agentic model for coding and long-horizon workflows. It supports image input, structured output, configurable reasoning, a 262K-token context window, and up to 235K output tokens.
Nex N2.5 Pro
nex-agi/nex-n2.5-proNex N2.5 Pro is an open-source multimodal agentic model for coding, tool use, and long-horizon workflows. It supports image input, structured output, configurable reasoning, a 262K-token context window, and up to 235K output tokens.