Self-hosted enterprise AI
Run a capable multimodal and agentic model in controlled infrastructure under a permissive Apache 2.0 license.
Independent tool overview
Mistral Small 4 is a 119B-parameter open-weight mixture-of-experts model that unifies fast instruction following, configurable reasoning, coding, vision and agentic tool use.
Visit the official Mistral 4 Small site ↗
Overview
Mistral Small 4 is a general-purpose hybrid model that consolidates capabilities previously split across Mistral's instruction, Magistral reasoning and Devstral coding families. A request can use fast response mode or spend additional compute in reasoning mode.
Its mixture-of-experts architecture contains 119 billion total parameters but activates about 6.5 billion per token. The model accepts text and images, supports a 256K-token context window and can return text, function calls and structured JSON.
Developers can use the managed Mistral API or download the official Apache 2.0 weights for their own infrastructure. The open-weight option offers deployment control, but the full model still requires substantial memory and serving expertise.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Run a capable multimodal and agentic model in controlled infrastructure under a permissive Apache 2.0 license.
Use one model for fast instruction following and higher-effort reasoning instead of routing to separate model families.
Build applications that need function calling, JSON output, long context and strong system-prompt adherence.
Capabilities
Switches between fast instruction mode and configurable reasoning effort within the same model.
Activates only 6.5B of 119B parameters per token to improve serving efficiency relative to a dense model of similar total size.
Analyzes images together with text for document, interface and visual-understanding tasks.
Supports native function calling, agent and conversation APIs, built-in tools and structured outputs.
Provides a 256K-token window for large documents, repositories and multi-step workflows.
Available through Mistral's hosted API or as downloadable FP8, NVFP4 and community-compatible checkpoints.
Process
Step 1
Use the Mistral API for fast integration or evaluate hardware, quantization and operations before self-hosting.
Step 2
Call mistral-small-2603 through Mistral's APIs or load the official Mistral-Small-4-119B-2603 weights.
Step 3
Set system instructions, reasoning effort, tools and the desired response format for the workload.
Step 4
Measure accuracy, latency, token use and memory requirements with representative prompts and images.
Step 5
Validate structured results, constrain tool permissions and monitor model or quantization changes after deployment.
Cost
The hosted Mistral API charges separately for input and output tokens. Official weights are available under Apache 2.0 without a model license fee, but self-hosting adds compute, storage and operations costs.
$0.15 input / $0.60 output per 1M tokens
Managed access using the mistral-small-2603 model ID.
No model license fee
Downloadable Apache 2.0 model weights for private or custom deployment.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
A family of smaller open models for deployments where hardware footprint matters more than Mistral Small 4's unified capability set.
Explore Qwen3.5 Small →Coding
A broad open-model ecosystem with extensive tooling and deployment support.
Explore Meta Llama Models →Consumer
A managed frontier model alternative for teams prioritizing hosted reasoning performance over open weights.
Explore Gemini 3.1 Pro →Questions
Mistral Small 4 is a 119B-parameter mixture-of-experts model that combines instruction following, reasoning, coding, vision and agent capabilities.
Mistral publishes the model weights under the permissive Apache 2.0 license. Open-weight is the most precise description because training data and the full training pipeline are not supplied.
The model has 119B total parameters and activates about 6.5B per token through its mixture-of-experts architecture.
Mistral lists $0.15 per million input tokens and $0.60 per million output tokens for mistral-small-2603.
Yes. It accepts text and image inputs and returns text output.
The weights can be self-hosted, and the model card lists vLLM, llama.cpp, LM Studio, SGLang and Transformers support. Hardware requirements remain significant for the full 119B model.
Bottom line
Mistral Small 4 is an unusually flexible open-weight model for teams that want reasoning, coding, vision and agent functions without maintaining several specialized models. The managed API is the simplest route; self-hosting makes sense when control or data locality justifies the hardware and operational burden.
Visit Mistral 4 Small website ↗
Nemotron 3 Super - NVIDIA's open 120B reasoning model with 1M token context for agentic workflows

GPT-5.4 mini & nano - OpenAI's fast, cheap small models built for coding agents and subagent workflows

Copilot Cowork - Microsoft's Anthropic-powered AI for running multi-step tasks across M365 apps

GLM-5-Turbo - Z AI's high-speed agentic model built specifically for OpenClaw

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.