Ling 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.
Added Sep 9, 2026
Model weightsContext Window
262.1K
Max Output
32.8K
Avg output tokens (7d)
1.8K tokens
Input Price (Auto)
$0.060/1M
Output Price (Auto)
$0.18/1M
Cache Read (Auto)
$0.012/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
24.6
Coding Index
57.0
Agentic Index
28.7
Agentic work
AutomationBench-AA
Workflow automation with guardrail penalties
15.7%
Better than 39% of models compared
AA-Briefcase
Agentic knowledge work (Elo)
984 Elo
Better than 48% of models compared
GDPval-AA v2
Economically valuable tasks (Elo)
1150 Elo
Better than 59% of models compared
Document reasoning
GDP.pdf
Professional PDF reasoning: all-pass rate
8.4%
Better than 35% of models compared
AA-LCR v1.1
Long context reasoning with updated grading
78.3%
Better than 81% of models compared
Reasoning
HLE
Humanity's Last Exam
22.0%
Better than 70% of models compared
CritPt
Research-level physics reasoning
2.0%
Coding
Terminal-Bench v4.0
Practical coding and terminal tasks
0.0%
Better than 16% of models compared
SciCode
Python programming for scientific computing
44.2%
Better than 37% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
14.3%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
22.0%
Legacy benchmarks
GPQA Diamond (legacy)
Graduate-level scientific reasoning
86.2%
Better than 81% of models compared
AA-LCR (unversioned / legacy)
Long context reasoning evaluation
78.3%
Better than 81% of models compared
GDPval-AA (unversioned / legacy)
Economically valuable tasks
32.5%
Last updated Oct 2, 2026
Artificial AnalysisProviders
Provider information for this model’s automatic routing. These routes cannot be selected individually.
Loading provider options…
Related text models
Compare Ling 3.0 Flash VL with similar models from the same provider or model family.
Ling 3.0 Flash
inclusionai/ling-3.0-flashLing-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
Ling 3.0 Flash Thinking
inclusionai/ling-3.0-flash:thinkingLing-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.
MiMo V2.6 Flash
xiaomi/mimo-v2.6-flashMiMo V2.6 Flash is Xiaomi's native omnimodal 309B-parameter mixture-of-experts model, activating 15B parameters per token. It balances intelligence, efficiency, and cost for coding, general agents, visual tasks, and cybersecurity, with text, image, video, and audio understanding and a 1M-token context window.
GLM 5.3 Flash Cybersecurity
z-ai/glm-5.3-flash-cybersecurityGLM 5.3 Flash Cybersecurity is a cybersecurity-focused variant based on the uncensored model, with provider moderation for illegal activities. It supports always-on reasoning, image understanding, tool calling, and a 1,048,576-token context window.
Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flashQwen3.8 Omni Flash is a fast omni-modal reasoning model for understanding text, images, audio, and video. It is especially suited to meeting summaries, transcripts and subtitles, speaker-aware audiovisual analysis, and long-form content review. It returns text and supports a nearly one-million-token context window, tool calling, and structured output.
DeepSeek V4.1 Flash TEE
TEE/deepseek-v4.1-flashDeepSeek V4.1 Flash supports text and image input, reasoning, tool calling, and structured output with a 1M-token context window. This route runs through Tinfoil attested inference inside a Trusted Execution Environment.