StepFun's most capable open-source reasoning model with visible reasoning traces. Built on a sparse Mixture-of-Experts architecture with 196B total parameters and only 11B active per token, it achieves frontier-level performance in math, logic, and agentic coding while reaching up to 350 tokens/sec. Supports 256K context. NOTE: This model runs via StepFun, which may log and train on your prompts.
Added Feb 2, 2026
Model weightsContext Window
256.0K
Max Output
256.0K
Avg output tokens (7d)
1.0K tokens
Input Price (Auto)
$0.10/1M
Output Price (Auto)
$0.30/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
16.6
Agentic work
T²-Bench Telecom (legacy)
Legacy fallback · Conversational AI agents in dual-control scenarios
94.4%
Better than 92% of models compared
Document reasoning
AA-LCR v1.1
Long context reasoning with updated grading
50.0%
Better than 44% of models compared
Reasoning
HLE
Humanity's Last Exam
21.1%
Better than 69% of models compared
IFBench
Instruction-following benchmark
64.6%
Better than 76% of models compared
CritPt
Research-level physics reasoning
2.5%
Coding
Terminal-Bench Hard (legacy)
Legacy fallback · Agentic coding and terminal use
27.3%
Better than 69% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
23.6%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
85.7%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
83.1%
Better than 73% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
50.0%
Better than 44% of models compared
Last updated Oct 1, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare Step 3.5 Flash with similar models from the same provider or model family.
Step 3.5 Flash 2603
stepfun-ai/step-3.5-flash-2603Step 3.5 Flash 2603 is optimized for high-frequency agentic and coding workflows with improved token efficiency and faster reasoning. NOTE: This model runs via StepFun, which may log and train on your prompts.
Step 3.7 Flash Thinking
stepfun/step-3.7-flash:thinkingStep 3.7 Flash Thinking is StepFun's high-efficiency multimodal MoE model with visible reasoning enabled for deeper agentic coding, long-context reasoning, tool use, and native image/video understanding. ⚠️ Note: This model routes through StepFun, so privacy and logging guarantees may be limited.
Step 5 Preview
stepfun/step-5-previewStep 5 Preview is StepFun's 600B sparse MoE frontier model for production-scale agents, activating 27B parameters per token. It is built for software engineering, long-horizon tool use, research, professional knowledge work, and finance, with native text, image, and video understanding and a 1M-token context window. ⚠️ Note: This model routes through StepFun, so privacy and logging guarantees may be limited.
MiMo V2.6 Flash
xiaomi/mimo-v2.6-flashMiMo V2.6 Flash is Xiaomi's native omnimodal 309B-parameter mixture-of-experts model, activating 15B parameters per token. It balances intelligence, efficiency, and cost for coding, general agents, visual tasks, and cybersecurity, with text, image, video, and audio understanding and a 1M-token context window.
GLM 5.3 Flash Cybersecurity
z-ai/glm-5.3-flash-cybersecurityGLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.
Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flashQwen3.8 Omni Flash is a fast omni-modal reasoning model for understanding text, images, audio, and video. It is especially suited to meeting summaries, transcripts and subtitles, speaker-aware audiovisual analysis, and long-form content review. It returns text and supports a nearly one-million-token context window, tool calling, and structured output.