The Rundown AI homepage

Independent tool overview

Mistral 4 Small at a glance

Mistral Small 4 is a 119B-parameter open-weight mixture-of-experts model that unifies fast instruction following, configurable reasoning, coding, vision and agentic tool use.

Visit the official Mistral 4 Small site ↗
Mistral 4 Small product preview
Model size
119B total, 6.5B active
Architecture
128-expert MoE, 4 active
Context window
256K tokens
Modalities
Text and image input; text output
License
Apache 2.0

Overview

What Mistral 4 Small is

Mistral Small 4 is a general-purpose hybrid model that consolidates capabilities previously split across Mistral's instruction, Magistral reasoning and Devstral coding families. A request can use fast response mode or spend additional compute in reasoning mode.

Its mixture-of-experts architecture contains 119 billion total parameters but activates about 6.5 billion per token. The model accepts text and images, supports a 256K-token context window and can return text, function calls and structured JSON.

Developers can use the managed Mistral API or download the official Apache 2.0 weights for their own infrastructure. The open-weight option offers deployment control, but the full model still requires substantial memory and serving expertise.

Use cases

Who Mistral 4 Small is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Self-hosted enterprise AI

Run a capable multimodal and agentic model in controlled infrastructure under a permissive Apache 2.0 license.

Mixed reasoning and chat workloads

Use one model for fast instruction following and higher-effort reasoning instead of routing to separate model families.

Coding and tool-using agents

Build applications that need function calling, JSON output, long context and strong system-prompt adherence.

Capabilities

Core Mistral 4 Small features

1

Hybrid response modes

Switches between fast instruction mode and configurable reasoning effort within the same model.

2

Sparse MoE architecture

Activates only 6.5B of 119B parameters per token to improve serving efficiency relative to a dense model of similar total size.

3

Vision input

Analyzes images together with text for document, interface and visual-understanding tasks.

4

Agent capabilities

Supports native function calling, agent and conversation APIs, built-in tools and structured outputs.

5

Long context

Provides a 256K-token window for large documents, repositories and multi-step workflows.

6

Flexible deployment

Available through Mistral's hosted API or as downloadable FP8, NVFP4 and community-compatible checkpoints.

Process

How the Mistral 4 Small workflow works

  1. Step 1

    Choose hosted or self-managed

    Use the Mistral API for fast integration or evaluate hardware, quantization and operations before self-hosting.

  2. Step 2

    Select the model ID

    Call mistral-small-2603 through Mistral's APIs or load the official Mistral-Small-4-119B-2603 weights.

  3. Step 3

    Configure task behavior

    Set system instructions, reasoning effort, tools and the desired response format for the workload.

  4. Step 4

    Test both quality and serving cost

    Measure accuracy, latency, token use and memory requirements with representative prompts and images.

  5. Step 5

    Add production controls

    Validate structured results, constrain tool permissions and monitor model or quantization changes after deployment.

Cost

Mistral 4 Small pricing and free plan

The hosted Mistral API charges separately for input and output tokens. Official weights are available under Apache 2.0 without a model license fee, but self-hosting adds compute, storage and operations costs.

Mistral API

$0.15 input / $0.60 output per 1M tokens

Managed access using the mistral-small-2603 model ID.

  • Pay for token usage
  • 256K context
  • Hosted chat, agents and tool features

Self-hosted weights

No model license fee

Downloadable Apache 2.0 model weights for private or custom deployment.

  • Infrastructure costs apply
  • FP8 and NVFP4 checkpoints available
  • vLLM is Mistral's recommended production server

Pricing checked . Check current pricing at the source ↗

Assessment

Mistral 4 Small strengths and limitations

Where it stands out

  • Combines instruction, reasoning, coding and vision capabilities in one model.
  • Apache 2.0 weights support commercial self-hosting and customization.
  • Only a fraction of total parameters activate per token.
  • Strong developer feature set including function calling and structured output.
  • Competitive hosted token pricing for a model with long context and multimodal input.

What to consider

  • The 119B checkpoint is still large and requires substantial GPU memory even with sparse activation.
  • Self-hosting shifts reliability, scaling, security and optimization work to the operator.
  • The model accepts images but produces text rather than generating visual media.
  • Long-context availability does not guarantee equally strong retrieval across every 256K-token prompt.
  • Quantization can improve throughput and memory use while reducing quality, especially on long context.

Compare

Mistral 4 Small alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Qwen3.5 Small

A family of smaller open models for deployments where hardware footprint matters more than Mistral Small 4's unified capability set.

Explore Qwen3.5 Small

Consumer

Gemini 3.1 Pro

A managed frontier model alternative for teams prioritizing hosted reasoning performance over open weights.

Explore Gemini 3.1 Pro

Questions

Mistral 4 Small FAQs

What is Mistral Small 4?

Mistral Small 4 is a 119B-parameter mixture-of-experts model that combines instruction following, reasoning, coding, vision and agent capabilities.

Is Mistral Small 4 open source?

Mistral publishes the model weights under the permissive Apache 2.0 license. Open-weight is the most precise description because training data and the full training pipeline are not supplied.

How many parameters does Mistral Small 4 use?

The model has 119B total parameters and activates about 6.5B per token through its mixture-of-experts architecture.

How much does the Mistral Small 4 API cost?

Mistral lists $0.15 per million input tokens and $0.60 per million output tokens for mistral-small-2603.

Can Mistral Small 4 analyze images?

Yes. It accepts text and image inputs and returns text output.

Can I run Mistral Small 4 locally?

The weights can be self-hosted, and the model card lists vLLM, llama.cpp, LM Studio, SGLang and Transformers support. Hardware requirements remain significant for the full 119B model.

Bottom line

Our Mistral 4 Small verdict

Mistral Small 4 is an unusually flexible open-weight model for teams that want reasoning, coding, vision and agent functions without maintaining several specialized models. The managed API is the simplest route; self-hosting makes sense when control or data locality justifies the hardware and operational burden.

Visit Mistral 4 Small website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.