Browse all Arcee Ai text models
Provider logo

Trinity Large Thinking

arcee-ai/trinity-large-thinking
Back
Provider logo

Trinity Large Thinking

arcee-ai/trinity-large-thinking
Back

Open source Arcee reasoning model with a 262K context window, 80K max output, and native reasoning and tool support for agentic workloads.

Added Apr 1, 2026

Model weights

Context Window

262.1K

Max Output

80.0K

Avg output tokens (7d)

242 tokens

19%

Input Price (Auto)

$0.25/1M

Output Price (Auto)

$0.90/1M

Capabilities

Benchmarks

Performance metrics and benchmarks

Sourced from Artificial Analysis.

Intelligence Index

10.8

Better than 42% of models compared

Coding Index

25.8

Better than 32% of models compared

Agentic Index

1.0

Better than 16% of models compared

Agentic work

AutomationBench-AA

Workflow automation with guardrail penalties

1.3%

Better than 16% of models compared

AA-Briefcase

Agentic knowledge work (Elo)

255 Elo

Better than 15% of models compared

GDPval-AA v2

Economically valuable tasks (Elo)

324 Elo

Better than 19% of models compared

Document reasoning

GDP.pdf

Professional PDF reasoning: all-pass rate

1.2%

Better than 9% of models compared

AA-LCR v1.1

Long context reasoning with updated grading

38.0%

Better than 35% of models compared

Reasoning

HLE

Humanity's Last Exam

15.8%

Better than 63% of models compared

IFBench

Instruction-following benchmark

56.3%

Better than 67% of models compared

CritPt

Research-level physics reasoning

0.9%

Coding

Terminal-Bench v4.0

Practical coding and terminal tasks

0.5%

Better than 35% of models compared

SciCode

Python programming for scientific computing

40.6%

Better than 29% of models compared

Knowledge

AA-Omniscience Accuracy

Proportion of correctly answered questions

22.5%

AA-Omniscience Hallucination Rate

Rate of incorrect answers among non-correct responses

85.9%

Legacy benchmarks

GPQA Diamond (legacy)

Graduate-level scientific reasoning

75.2%

Better than 57% of models compared

Terminal-Bench Hard (legacy)

Agentic coding and terminal use

22.7%

Better than 62% of models compared

T²-Bench Telecom (legacy)

Conversational AI agents in dual-control scenarios

90.1%

Better than 84% of models compared

AA-LCR (unversioned / legacy)

Long context reasoning evaluation

38.0%

Better than 35% of models compared

GDPval-AA (unversioned / legacy)

Economically valuable tasks

0.0%

Last updated Oct 2, 2026

Artificial Analysis

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare Trinity Large Thinking with similar models from the same provider or model family.

Abliterated Model Large V2

abliteration-ai/abliterated-model-large-v2

Abliteration.ai's default unrestricted large text reasoning model is derived from GLM-5.3 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.

Abliterated Model Large

abliteration-ai/abliterated-model-large

Abliteration.ai's unrestricted large text reasoning model is derived from GLM-5.2 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.

Hermes 4 Large

nousresearch/hermes-4-405b

Advanced reasoning model built on Llama-3.1-405B with hybrid thinking modes. Features internal deliberation capabilities, excels at math, code, STEM, and logical reasoning while supporting structured outputs with improved steerability and neutral alignment.

Llama 3.3 70B Wayfarer

LatitudeGames/Wayfarer-Large-70B-Llama-3.3

Llama 3.3 70B Wayfarer is a fine-tuned version of Llama 3.3 70B, trained on a diverse set of creative writing and RP datasets with a focus on variety and deduplication. This model is designed to be highly creative and non-repetitive by making sure no two entries in the dataset have repeated characters or situations, which makes sure the model does not latch on to a certain personality and be capable of understanding and acting appropriately to any characters or situations.

Mistral Large 2411

mistralai/mistral-large

Upgrade to Mistral's flagship model. It is fluent in English, French, Spanish, German, and Italian, with high grammatical accuracy, with a long context window.

Hermes 4 Large (Thinking)

nousresearch/hermes-4-405b:thinking

Hermes 4 Large with thinking enabled. Streams visible reasoning before the final answer and supports structured outputs.