Large codebase work
Analyze repositories and maintain longer working histories without immediately compressing the project context.
Independent tool overview
MiniMax M3 is a large open-weight multimodal model for coding and agentic work, combining a one-million-token context window, image and video understanding, computer use and optional reasoning.
Visit the official MiniMax M3 site ↗
Overview
MiniMax M3 is designed for long-running coding and computer-based work rather than short chat alone. Its MiniMax Sparse Attention architecture supports up to one million tokens of context, and the model was trained natively across text, images and video.
Users can access M3 through MiniMax Code, monthly Token Plans, a pay-as-you-go API or downloadable weights. The standard model has roughly 428 billion total parameters with about 23 billion activated per token, so self-hosting the full release is possible but requires substantial infrastructure.
“Open weight” does not mean unrestricted open source. The MiniMax Community License allows broad use but includes notice and authorization requirements for commercial products, including prior written authorization when annual product or service revenue exceeds $20 million.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Analyze repositories and maintain longer working histories without immediately compressing the project context.
Plan, invoke tools, write code, run checks and revise work across multi-step software tasks.
Combine screenshots, video or visual interfaces with text instructions in desktop and office workflows.
Evaluate downloadable weights when control over hosting and data flow matters enough to justify significant infrastructure.
Capabilities
MiniMax Sparse Attention is designed to reduce long-context compute while retaining information across extended coding and agent sessions.
The model accepts text, images and video, supporting visual analysis and computer-use tasks alongside code and language work.
API and model settings can keep thinking enabled, choose it adaptively or disable it for lower-latency responses.
MiniMax's official coding agent pairs M3 with long-running, multi-agent workflows and computer use.
The hosted API and common local serving stacks can expose chat-completions-style interfaces for existing agent tools.
Official weights are available on Hugging Face with documented Transformers, vLLM, SGLang and other serving paths.
Process
Step 1
Use MiniMax Code or the hosted API for quick adoption; consider open weights only when deployment control justifies the hardware and operational burden.
Step 2
Start with adaptive thinking, then test enabled mode for harder tasks and disabled mode for completions or latency-sensitive chat.
Step 3
Give the agent the repository, visual inputs and narrow permissions it needs, with clear completion criteria.
Step 4
Monitor long-context growth, cache usage, rolling subscription windows and any production API spend.
Step 5
Measure task success, latency, tool reliability, security and human correction rates against alternative models.
Cost
MiniMax sells monthly Token Plans for interactive developer use and a separate pay-as-you-go API for production. Current M3 API rates are discounted 50%; inputs above 512K tokens cost twice the shorter-context rate.
$20/month
Entry plan for individual coding and daily agent workflows.
$50/month
Higher-quota plan for daily professional work.
$120/month
Highest interactive quota for heavy individual or team workflows.
$0.30 input / $1.20 output per 1M tokens
Current discounted pay-as-you-go rate for standard-context production requests.
$0.60 input / $2.40 output per 1M tokens
Long-context rate for requests between 512K and 1M input tokens.
No model download fee
Self-host the official M3 weights under the MiniMax Community License.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consulting
Compare Kimi K3 for another current open model aimed at frontier performance and lower-cost deployment.
Explore Kimi K3 →Consumer
Compare GLM 5.3 for open-weight coding and agentic work in Z AI's ecosystem.
Explore GLM-5.3 →Consumer
Choose GPT-5.5 when OpenAI's managed tool stack and production model support matter more than downloadable weights.
Explore GPT 5.5 →Questions
MiniMax M3 is a natively multimodal model for coding and agentic work. It supports up to one million tokens of context, image and video input, computer use, configurable thinking and downloadable weights.
MiniMax describes M3 as open weight. The files use the MiniMax Community License, which is not an unrestricted standard open-source license and includes commercial notice and authorization conditions.
The official model card lists about 428 billion total parameters and roughly 23 billion activated parameters per token.
At the current 50% discount, requests up to 512K input tokens cost $0.30 per million input tokens and $1.20 per million output tokens. Requests above 512K cost $0.60 input and $2.40 output.
MiniMax advertises up to one million tokens. Its model page says the hosted API guarantees at least 512K, with the 512K-to-1M range billed at the long-context rate.
Yes. M3 was trained natively on mixed modalities and accepts image and video input along with text.
Yes, official weights and serving instructions are available. Because the full model is roughly 428B parameters, practical local deployment requires substantial accelerator memory or a specialized quantized setup.
The model supports enabled thinking, adaptive thinking and disabled thinking. Adaptive is a reasonable starting point; teams should benchmark the quality and latency tradeoff for their workload.
Bottom line
MiniMax M3 is notable for putting long context, native multimodality and serious agent capabilities into a downloadable model while also offering inexpensive hosted access. The API is the practical route for most teams. Self-hosting makes sense only when control or data requirements outweigh the cost of serving a 428B-parameter model, and the community license should be reviewed before commercial use.
Visit MiniMax M3 website ↗
Claude Opus 4.8 - Anthropic's new top model with improvements to reliability, coding, and agentic flows.

Nemotron 3 Ultra - Nvidia’s open 550B reasoning model for agents

Qwen 3.7 Max- Alibaba's new flagship model for long-horizon agentic tasks

Claude Fable - Anthropic's new frontier Mythos-class model with top performance across benchmarks

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.