Long-horizon coding agents
Run repository-scale work that may require many edits, tests, terminal calls, and recovery steps.
Independent tool overview
MiMo-V2.5-Pro is Xiaomi's MIT-licensed, text-only Mixture-of-Experts model for long-horizon coding and agent tasks, with 1.02 trillion total parameters, 42 billion active parameters, and a 1-million-token context window.
Visit the official MiMo-V2.5-Pro site ↗
Overview
MiMo-V2.5-Pro is Xiaomi MiMo's flagship open-weight language model for demanding software engineering, reasoning, and tool-using agents. It is available through Xiaomi's pay-as-you-go API, Token Plan subscriptions, AI Studio, and downloadable FP8 weights.
The model uses a 1.02-trillion-parameter Mixture-of-Experts architecture while activating 42 billion parameters per token. Hybrid sliding-window and global attention reduce long-context cache pressure, and three Multi-Token Prediction modules are designed to increase decoding throughput.
This Pro variant is text-only despite its 1-million-token context. Xiaomi's separate MiMo-V2.5 model supports text and image input. Teams integrating Pro into coding agents also need a compatible harness that preserves the model's reasoning content across tool-call turns.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Run repository-scale work that may require many edits, tests, terminal calls, and recovery steps.
Provide extensive source, documentation, and task history in one long context.
Pair the model with a capable harness for search, file editing, commands, testing, and structured tools.
Use low published token prices for high-volume text reasoning and code generation.
Run MIT-licensed weights on owned infrastructure when the organization can support a trillion-parameter MoE.
Capabilities
Supports up to 1,048,576 input tokens for large repositories and long-running agent histories.
Post-training targets coherent tool use and instruction following across extended, multi-step tasks.
Designed for repository understanding, implementation, review, structured artifacts, and software-engineering workflows.
Interleaves local sliding-window and global attention at a 6:1 ratio to reduce KV-cache requirements.
Uses three native MTP modules intended to improve inference throughput.
Xiaomi provides a compatible endpoint for pay-as-you-go access and region-specific Token Plans.
Publishes instruct and base checkpoints under the permissive MIT license.
Provides official deployment guidance for SGLang and vLLM, including reasoning and tool-call parsers.
Process
Step 1
Compare data requirements and sustained volume against the infrastructure needed for 1.02T FP8 weights.
Step 2
Use a coding agent that supports Xiaomi's endpoint, thinking mode, long context, and tool-call replay rules.
Step 3
Set the correct model ID, context, output cap, parser, and provider-specific reasoning fields.
Step 4
Restrict file, terminal, network, deployment, and credential access to what the task actually needs.
Step 5
Test repository navigation, edits, tests, and recovery on a representative job before long autonomous runs.
Step 6
Replay the actual reasoning_content and assistant content required by the API across tool turns.
Step 7
Track pass rate, tokens, wall time, retries, tool errors, regressions, and human review—not benchmark scores alone.
Cost
Xiaomi offers very low pay-as-you-go token prices plus fixed Token Plan subscriptions. Pay-as-you-go separates cached input, uncached input, and output. Token Plans use dedicated regional endpoints and quota-based access; confirm the current included Credits in Subscription Management.
$0.0036 cached input / $0.435 uncached input / $0.87 output per 1M tokens
Usage-based access through Xiaomi's OpenAI-compatible API.
$6/month or $5.28/month effective annually
Entry subscription for quota-based model access.
$16/month or $14.08/month effective annually
Mid-volume subscription for individual development workflows.
$50 or $100 monthly; $44 or $88/month effective annually
Higher-quota subscriptions for heavier agent and coding use.
Free MIT license; infrastructure costs apply
Download and serve the FP8 model with compatible inference software.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
Choose Devstral 2 for Mistral's coding-specific open model family and a different deployment-to-capability tradeoff.
Explore Devstral 2 →Coding
Evaluate Kimi K2.7 Code for another current open-weight model focused on coding agents and leaner reasoning.
Explore Kimi K2.7 Code →Business Operations
Consider DeepSeek for a broader family of low-cost reasoning and coding models with its own API and open weights.
Explore DeepSeek →Consumer
Use Trinity Large Thinking for another permissively licensed long-context agent model with different routing and hosting options.
Explore Trinity-Large-Thinking →Questions
It is Xiaomi's flagship text-only open-weight MoE model for long-horizon coding, reasoning, and tool-using agents.
Xiaomi publishes the weights, tokenizer, and model card under the permissive MIT license.
It has 1.02 trillion total parameters and activates 42 billion parameters per token. The official instruct checkpoint uses FP8 mixed precision.
The instruct model supports up to 1,048,576 tokens. Xiaomi's current agent integration guide lists a 131,072-token maximum output.
No. The Pro integration is text-only. Xiaomi's separate MiMo-V2.5 model supports text and image input.
Pay-as-you-go pricing is $0.0036 per million cached input tokens, $0.435 per million uncached input tokens, and $0.87 per million output tokens.
Yes, with MIT-licensed weights and official SGLang and vLLM guidance, but the trillion-parameter FP8 model requires serious multi-GPU infrastructure.
A compatible harness must preserve the model's actual reasoning_content and assistant content in conversation history across tool-call turns.
Bottom line
MiMo-V2.5-Pro is compelling for text-only coding agents that need very long context and low hosted inference cost. The MIT license and detailed deployment material make it more open than many flagship alternatives. The practical decision is less about headline benchmarks and more about the harness: verify tool-call replay, long-context behavior, repository pass rate, and safe permissions. Use Xiaomi's API for an inexpensive proof before considering the formidable cost of self-hosting a 1.02T FP8 model.
Visit MiMo-V2.5-Pro website ↗
Nemotron 3 Nano Omni - NVIDIA's new open model combining vision, audio, and text
.jpeg)
Grok 4.3 - xAI’s latest model release with strong cost efficiency, instruction following, and domain-specific performance

GPT-5.5 - OpenAI's new frontier flagship model for agentic coding and computer use

ERNIE 5.1 - Baidu's new foundation model that ranks No. 4 on Arena search and claims 94% lower training costs

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.