Devstral 2 123B is a 123 billion parameter model from Mistral AI optimized for coding and development tasks. Features advanced reasoning capabilities for software engineering workflows.
Added Dec 9, 2025
Model weightsContext Window
262.1K
Max Output
65.5K
Avg output tokens (7d)
475 tokens
Input Price (Auto)
$0.40/1M
Output Price (Auto)
$1.40/1M
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
8.6
Coding Index
31.3
Agentic Index
2.3
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
3.1%
Better than 24% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
479 Elo
Better than 23% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
531 Elo
Better than 28% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
2.4%
Better than 14% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
32.3%
Better than 32% of models compared
Reasoning
HLE
Humanity's Last Exam
3.6%
Better than 6% of models compared
IFBench
Instruction-following benchmark
38.1%
Better than 34% of models compared
CritPt
Research-level physics reasoning
0.0%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 16% of models compared
SciCode
Python programming for scientific computing
32.8%
Better than 12% of models compared
LiveCodeBench
Contamination-free coding benchmark
44.8%
Better than 51% of models compared
Math
AIME 2025
American Invitational Mathematics Examination 2025
36.7%
Better than 37% of models compared
Knowledge
MMLU-Pro
Professional and academic subject knowledge
76.2%
Better than 53% of models compared
AA-Omniscience Accuracy
Proportion of correctly answered questions
20.8%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.3%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
59.4%
Better than 34% of models compared
Terminal-Bench Hard (legacy)
Agentic coding and terminal use
18.9%
Better than 59% of models compared
T²-Bench Telecom (legacy)
Conversational AI agents in dual-control scenarios
24.9%
Better than 28% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
32.3%
Better than 32% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
1.6%
Last updated Oct 1, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Devstral 2 123B with similar models from the same provider or model family.
Ministral 3 14B
mistralai/ministral-14b-instruct-2512Ministral 3 14B is a balanced model in the Ministral 3 family, designed for edge deployment. A powerful, efficient language model with vision capabilities, fine-tuned for instruction tasks. Features multilingual support, strong system prompt adherence, and native function calling. Apache 2.0 licensed.
Mistral Small 3 24B (2501)
mistralai/mistral-small-24b-instruct-2501Mistral Small 3 24B (2501) hosted by IONOS in Berlin, Germany. Zero data retention.
Mixtral 8x22B
mistralai/mixtral-8x22b-instruct-v0.1Mixtral 8x22B is a powerful sparse Mixture of Experts (MoE) model with 141B total parameters and 39B active per token. Features a 64K context window, exceptional math performance, and cost-efficient inference. Supports English, French, German, Spanish, and Italian. Apache 2.0 licensed.
Ministral 14B
mistralai/ministral-14b-2512Ministral 14B is a powerful 14B parameter model from Mistral AI with vision capabilities, offering frontier performance in a compact size.
Ministral 3B
mistralai/ministral-3b-2512Ministral 3B is a tiny, efficient 3B parameter model from Mistral AI with vision capabilities, designed for edge deployment.
Ministral 8B
mistralai/ministral-8b-2512Ministral 8B is an efficient 8B parameter model from Mistral AI with vision capabilities, designed for edge deployment.