Long-running agents
Maintain reasoning and tool state across multi-step workflows where a model must plan, act, inspect results, and continue.
Independent tool overview
Trinity-Large-Thinking is Arcee AI's open-weight 398B sparse MoE reasoning model for tool calling, multi-step planning, coding, and long-running agent loops.
Visit the official Trinity-Large-Thinking site ↗
Overview
Trinity-Large-Thinking is the reasoning-optimized checkpoint in Arcee AI's Trinity Large family. It has roughly 398 billion total parameters but activates about 13 billion per token through 4-of-256 expert routing, aiming to combine frontier-scale capacity with more efficient inference than a similarly sized dense model.
The model is post-trained with extended chain-of-thought and agentic reinforcement learning for multi-step planning, tool calling, coding, and customer-service-style workflows. It emits explicit reasoning in think blocks or a reasoning_content field, and Arcee warns that this reasoning must be preserved in message history for reliable multi-turn agent behavior.
Arcee's model card describes a 512K extended context window, while the current OpenRouter endpoint exposes 262,144 tokens. The downloadable weights use the permissive OpenMDW-1.1 license, but a 398B checkpoint remains operationally demanding even with quantized variants, so hosted API access is the practical starting point for most teams.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Maintain reasoning and tool state across multi-step workflows where a model must plan, act, inspect results, and continue.
Orchestrate structured lookups and actions across service, telecom, travel, and other process-driven workflows.
Power coding agents that need planning, command execution, repository work, and iterative verification.
Run or customize a frontier-scale reasoning model when the organization needs control over weights and infrastructure.
Evaluate a lower-priced hosted alternative to closed frontier reasoning models for high token volumes.
Capabilities
Routes each token through four experts plus one shared expert, activating about 13B parameters rather than the full model.
Produces extended thinking that can be returned separately from the final answer and preserved across turns.
Supports structured tool calls for agent loops, with documented parsers and OpenAI-compatible serving patterns.
The checkpoint is extended to 512K tokens, though hosted providers may expose a smaller window.
Arcee documents compatibility with OpenClaw and Hermes Agent, and recommends vLLM for agentic deployment.
Official releases include the main checkpoint plus FP8, W4A16, NVFP4, and GGUF variants.
OpenMDW-1.1 allows use, modification, redistribution, and commercialization subject to its notice and termination terms.
Process
Step 1
Prototype on Arcee Chat or OpenRouter before committing to the infrastructure required for self-hosting.
Step 2
Give each tool a clear schema, least-privilege permissions, deterministic validation, and an explicit failure path.
Step 3
Configure the provider or vLLM parser so reasoning and final content are returned in the expected fields.
Step 4
Append reasoning_content, content, tool calls, and tool results to message history for every agent step.
Step 5
Budget for thinking tokens and remove complete older turns when needed instead of stripping recent reasoning.
Step 6
Measure tool-selection accuracy, malformed calls, recovery behavior, task completion, latency, and total token cost.
Step 7
Require confirmation, policy checks, idempotency, and audit logs before tools create irreversible or high-impact changes.
Cost
The weights are downloadable without a model usage fee under OpenMDW-1.1, while infrastructure is separate. OpenRouter pricing varies by provider and routing; its current model page lists routed pricing around $0.22 per million input tokens, $0.85 per million output tokens, and $0.06 per million cached-read tokens.
$0.22 input / $0.85 output per 1M tokens
OpenRouter can route requests across available providers for price, speed, or reliability.
$0.25 input / $0.80 output per 1M tokens
Pin the request to Arcee AI's current hosted provider rate when consistency matters.
No model usage fee
Run the official or quantized checkpoints on infrastructure you control.
No public rate card
Arcee links to a browser chat for trying the model, but the model card does not publish a dedicated plan price.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Evaluate GLM-5.1 for another open model focused on long-horizon agentic coding.
Explore GLM-5.1 →Consumer
Consider MiniMax M2.7 when strong coding and agent benchmarks matter and a different hosted or open model stack is acceptable.
Explore Minimax M2.7 →Agents
Use Kimi K2.5 for an alternative open agent-focused model with its own tool-use and coding strengths.
Explore Kimi K2.5 →Consumer
Choose Claude Opus 4.6 when managed frontier performance and a larger commercial ecosystem matter more than open weights.
Explore Claude Opus 4.6 →Questions
It is Arcee AI's reasoning-optimized 398B sparse mixture-of-experts model for planning, coding, tool calling, and multi-step agent workflows.
The model has roughly 398B total parameters and activates about 13B per token through 4-of-256 expert routing plus a shared expert.
Arcee's model card lists a 512K extended context window. The current OpenRouter endpoint lists 262,144 tokens, so applications must use the provider's actual limit.
OpenRouter currently lists routed pricing around $0.22 per million input tokens, $0.85 per million output tokens, and $0.06 per million cached-read tokens. Exact provider rates can differ.
Yes. OpenMDW-1.1 is permissive and allows commercial use, subject to retaining the license and applicable origin notices on redistribution and its other terms.
Arcee says the model relies on its recent thinking to continue multi-step work. Omitting it can degrade agent behavior and cause malformed tool calls.
Yes. Arcee publishes full and quantized checkpoints with vLLM, SGLang, Transformers, and Docker guidance, but the model's total size requires serious infrastructure planning.
The official model card describes a text-generation model. Some API gateways may accept additional content formats, but do not assume native image or audio understanding without endpoint-specific documentation and tests.
Bottom line
Trinity-Large-Thinking is a credible open-weight option for teams building tool-heavy agents and willing to implement reasoning preservation correctly. Its low hosted price, sparse activation, permissive license, long context, and official quantizations are attractive. The main risks are operational: a huge checkpoint, provider-specific context limits, fast-growing reasoning history, and the need to secure every tool action. Benchmark it inside the real harness rather than choosing from leaderboard scores alone.
Visit Trinity-Large-Thinking website ↗
Qwen 3.5 Omni - Alibaba's native omnimodal model with text, image, audio, and video understanding across 113 languages

LFM2.5-350M - Liquid AI's 350M-parameter edge model built for tool use and on-device agents

Cohere Transcribe- SOTA, free open-source speech recognition model
Gemma 4 - Google's new open-weight reasoning and agentic model family across four sizes

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.