The Rundown AI homepage

Independent tool overview

MiniMax M3 at a glance

MiniMax M3 is a large open-weight multimodal model for coding and agentic work, combining a one-million-token context window, image and video understanding, computer use and optional reasoning.

Visit the official MiniMax M3 site ↗
MiniMax M3 product preview
Best for
Long-horizon coding and agent workflows
Context window
Up to 1M tokens; API minimum guaranteed at 512K
Modalities
Text, image and video input
Model size
~428B total / ~23B activated parameters
Access
MiniMax Code, API, Token Plan and open weights
License
MiniMax Community License

Overview

What MiniMax M3 is

MiniMax M3 is designed for long-running coding and computer-based work rather than short chat alone. Its MiniMax Sparse Attention architecture supports up to one million tokens of context, and the model was trained natively across text, images and video.

Users can access M3 through MiniMax Code, monthly Token Plans, a pay-as-you-go API or downloadable weights. The standard model has roughly 428 billion total parameters with about 23 billion activated per token, so self-hosting the full release is possible but requires substantial infrastructure.

“Open weight” does not mean unrestricted open source. The MiniMax Community License allows broad use but includes notice and authorization requirements for commercial products, including prior written authorization when annual product or service revenue exceeds $20 million.

Use cases

Who MiniMax M3 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Large codebase work

Analyze repositories and maintain longer working histories without immediately compressing the project context.

Agentic coding

Plan, invoke tools, write code, run checks and revise work across multi-step software tasks.

Multimodal computer use

Combine screenshots, video or visual interfaces with text instructions in desktop and office workflows.

Teams comparing open-weight frontier models

Evaluate downloadable weights when control over hosting and data flow matters enough to justify significant infrastructure.

Capabilities

Core MiniMax M3 features

1

One-million-token context

MiniMax Sparse Attention is designed to reduce long-context compute while retaining information across extended coding and agent sessions.

2

Native multimodality

The model accepts text, images and video, supporting visual analysis and computer-use tasks alongside code and language work.

3

Three reasoning modes

API and model settings can keep thinking enabled, choose it adaptively or disable it for lower-latency responses.

4

MiniMax Code

MiniMax's official coding agent pairs M3 with long-running, multi-agent workflows and computer use.

5

OpenAI-compatible access

The hosted API and common local serving stacks can expose chat-completions-style interfaces for existing agent tools.

6

Downloadable weights

Official weights are available on Hugging Face with documented Transformers, vLLM, SGLang and other serving paths.

Process

How the MiniMax M3 workflow works

  1. Step 1

    Choose hosted or self-managed access

    Use MiniMax Code or the hosted API for quick adoption; consider open weights only when deployment control justifies the hardware and operational burden.

  2. Step 2

    Set the reasoning mode

    Start with adaptive thinking, then test enabled mode for harder tasks and disabled mode for completions or latency-sensitive chat.

  3. Step 3

    Provide bounded tools and context

    Give the agent the repository, visual inputs and narrow permissions it needs, with clear completion criteria.

  4. Step 4

    Track quota and token behavior

    Monitor long-context growth, cache usage, rolling subscription windows and any production API spend.

  5. Step 5

    Evaluate before production

    Measure task success, latency, tool reliability, security and human correction rates against alternative models.

Cost

MiniMax M3 pricing and free plan

MiniMax sells monthly Token Plans for interactive developer use and a separate pay-as-you-go API for production. Current M3 API rates are discounted 50%; inputs above 512K tokens cost twice the shorter-context rate.

Plus Token Plan

$20/month

Entry plan for individual coding and daily agent workflows.

  • Approximately 1.7B M3 tokens per month
  • About 3–4 concurrent agents
  • Shared text, image and speech quota
  • Five-hour rolling and weekly usage windows

Max Token Plan

$50/month

Higher-quota plan for daily professional work.

  • Approximately 5.1B M3 tokens per month
  • About 4–5 concurrent agents
  • Three video generations per day
  • Five-hour rolling and weekly usage windows

Ultra Token Plan

$120/month

Highest interactive quota for heavy individual or team workflows.

  • Approximately 12.5B M3 tokens per month
  • About 6–7 concurrent agents
  • Five video generations per day
  • Five-hour rolling and weekly usage windows

M3 API up to 512K input

$0.30 input / $1.20 output per 1M tokens

Current discounted pay-as-you-go rate for standard-context production requests.

  • $0.06 per 1M cache-read tokens
  • Thinking modes share the same token rates
  • Production use is recommended on pay-as-you-go rather than Token Plans

M3 API above 512K input

$0.60 input / $2.40 output per 1M tokens

Long-context rate for requests between 512K and 1M input tokens.

  • $0.12 per 1M cache-read tokens
  • Applies to full-repository and ultra-long document workloads
  • Priority service is handled separately

Open weights

No model download fee

Self-host the official M3 weights under the MiniMax Community License.

  • Infrastructure and engineering costs are separate
  • Commercial notice or authorization requirements may apply
  • Full weights are large and operationally demanding

Pricing checked . Check current pricing at the source ↗

Assessment

MiniMax M3 strengths and limitations

Where it stands out

  • Combines long context, coding, multimodal input and computer use in one open-weight model
  • Hosted API pricing is low relative to many closed frontier models
  • Adaptive, enabled and disabled thinking modes support different latency needs
  • Official Token Plans work with MiniMax Code and common coding tools
  • Multiple documented local-serving options reduce integration lock-in

What to consider

  • The 428B-parameter weight set is expensive and complex to self-host
  • The MiniMax Community License includes commercial notice and revenue-based authorization conditions
  • Token Plan quotas use rolling and weekly windows and can tighten during peak traffic
  • MiniMax recommends pay-as-you-go rather than Token Plans for production workloads
  • Inputs above 512K tokens double the M3 API token rates
  • Official benchmark claims should be validated on the team's own tasks and agent harness
  • Computer-use and long-running agents require permission controls, monitoring and independent verification

Compare

MiniMax M3 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consulting

Kimi K3

Compare Kimi K3 for another current open model aimed at frontier performance and lower-cost deployment.

Explore Kimi K3

Consumer

GLM-5.3

Compare GLM 5.3 for open-weight coding and agentic work in Z AI's ecosystem.

Explore GLM-5.3

Consumer

GPT 5.5

Choose GPT-5.5 when OpenAI's managed tool stack and production model support matter more than downloadable weights.

Explore GPT 5.5

Questions

MiniMax M3 FAQs

What is MiniMax M3?

MiniMax M3 is a natively multimodal model for coding and agentic work. It supports up to one million tokens of context, image and video input, computer use, configurable thinking and downloadable weights.

Is MiniMax M3 open source?

MiniMax describes M3 as open weight. The files use the MiniMax Community License, which is not an unrestricted standard open-source license and includes commercial notice and authorization conditions.

How large is MiniMax M3?

The official model card lists about 428 billion total parameters and roughly 23 billion activated parameters per token.

How much does the MiniMax M3 API cost?

At the current 50% discount, requests up to 512K input tokens cost $0.30 per million input tokens and $1.20 per million output tokens. Requests above 512K cost $0.60 input and $2.40 output.

What is the MiniMax M3 context window?

MiniMax advertises up to one million tokens. Its model page says the hosted API guarantees at least 512K, with the 512K-to-1M range billed at the long-context rate.

Can MiniMax M3 understand images and video?

Yes. M3 was trained natively on mixed modalities and accepts image and video input along with text.

Can you run MiniMax M3 locally?

Yes, official weights and serving instructions are available. Because the full model is roughly 428B parameters, practical local deployment requires substantial accelerator memory or a specialized quantized setup.

What are MiniMax M3's thinking modes?

The model supports enabled thinking, adaptive thinking and disabled thinking. Adaptive is a reasonable starting point; teams should benchmark the quality and latency tradeoff for their workload.

Bottom line

Our MiniMax M3 verdict

MiniMax M3 is notable for putting long context, native multimodality and serious agent capabilities into a downloadable model while also offering inexpensive hosted access. The API is the practical route for most teams. Self-hosting makes sense only when control or data requirements outweigh the cost of serving a 428B-parameter model, and the community license should be reviewed before commercial use.

Visit MiniMax M3 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.