On-device applications
Evaluate E2B or E4B when latency, privacy, offline use, or mobile deployment matters more than maximum capability.
Independent tool overview
Gemma 4 is Google's five-model family of Apache 2.0-licensed open-weight models for reasoning, coding, agents, and multimodal understanding across devices from high-end phones to servers.
Visit the official Gemma 4 site ↗Overview
Gemma 4 includes E2B, E4B, 12B Unified, 26B A4B, and 31B models in pre-trained and instruction-tuned variants. The smaller E2B and E4B models target mobile and laptop use; 12B targets laptops, desktops, and small servers; 26B A4B uses a mixture-of-experts design that activates about 4B parameters per token; and 31B targets larger servers or clusters.
All five models accept text and image input and produce text. E2B, E4B, and 12B also accept audio, while video is handled as image frames. The family adds configurable thinking, native function calling and system prompts, long context, multilingual support, and dedicated draft models for speculative decoding. Those capabilities make Gemma 4 flexible, but deployment still requires model selection, hardware planning, prompt-template compliance, evaluation, safety controls, and application code around the weights.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Evaluate E2B or E4B when latency, privacy, offline use, or mobile deployment matters more than maximum capability.
Use E4B or 12B for local multimodal assistants, document tools, and developer workflows on capable consumer hardware.
Start with 26B A4B when the mixture-of-experts architecture offers the right quality, speed, and memory tradeoff.
Use native function-call generation for tools and workflows, with application-side schema validation, permissions, and execution.
Process documents, charts, screenshots, OCR, handwriting, images, short audio, and video frames into text.
Fine-tune or adapt downloadable weights when a domain-specific deployment justifies the data and evaluation work.
Capabilities
Covers efficient edge models, a unified 12B model, a sparse 26B A4B model, and a dense 31B model.
Supports reasoning modes that trade extra latency and tokens for more deliberate responses.
Provides 128K-token windows on E2B and E4B and 256K on 12B, 26B A4B, and 31B.
Handles OCR, documents and PDFs, charts, handwriting, user interfaces, object detection, and pointing tasks.
E2B, E4B, and 12B can transcribe or translate short speech inputs.
Processes videos as sequences of frames, with quality and compute shaped by sampling and visual-token settings.
Generates structured tool calls for an application to validate and execute.
Supports code generation, completion, correction, and agent-oriented development tasks.
Adds native support for the system role to control assistant behavior more cleanly.
Google reports pre-training across more than 140 languages and out-of-the-box support for 35 or more.
Lets developers choose visual-token budgets from 70 to 1,120 to balance detail and compute.
Dedicated multi-token-prediction draft models can accelerate inference without changing the target model's output quality.
Official weights are available through Hugging Face and Kaggle with documentation for common inference stacks.
Process
Step 1
List modalities, languages, latency, privacy, context, tool use, quality, and cost requirements using representative examples.
Step 2
Start with the smallest model that may pass the task, or Google's suggested 26B A4B general-purpose starting point for server-class use.
Step 3
Benchmark full or quantized weights on the actual device, inference library, batch size, and target context length.
Step 4
Use Gemma 4's current chat, thinking, modality-order, and function-calling formatting rather than an older Gemma template.
Step 5
Validate tool calls, restrict permissions, filter untrusted inputs, manage context, and keep generated reasoning out of normal conversation history as documented.
Step 6
Measure factuality, task success, latency, memory, safety, language quality, multimodal accuracy, and failure recovery on your own data.
Step 7
Track model revisions, dependency changes, drift, user-reported failures, hardware cost, and regressions after tuning or quantization.
Cost
Google publishes Gemma 4 weights under the Apache 2.0 license without a per-token model-license fee. Total cost depends on hardware or cloud compute, storage, bandwidth, inference software, engineering, observability, security, fine-tuning, and support. Hosted providers may charge their own rates.
No per-token license fee
Download and run the official models under Apache 2.0, while paying the operational cost of the chosen deployment.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Qwen3.5-Omni is a close alternative when broad text, image, audio, and video understanding across many languages is the priority.
Explore Qwen3.5-Omni →Consumer
Mistral 4 Small is a larger sparse open model to compare for reasoning, coding, vision, and server deployment.
Explore Mistral 4 Small →Coding
Meta's Llama family offers another widely supported downloadable-model ecosystem with different licensing and deployment tradeoffs.
Explore Meta Llama Models →Questions
Gemma 4 is Google's family of downloadable models for reasoning, coding, agentic tool use, and multimodal understanding across mobile, laptop, desktop, and server targets.
There are five core sizes: E2B, E4B, 12B Unified, 26B A4B, and 31B. The family includes pre-trained and instruction-tuned variants.
Google describes Gemma as an open-model family and publishes Gemma 4 weights under the Apache 2.0 license. Review that license and the distribution package for your use case.
Yes. E2B and E4B target mobile and laptop use, while larger variants need progressively more capable hardware. Actual feasibility depends on precision, runtime, context, batch size, and available memory.
Use the smallest model that might meet your quality target. Google's getting-started guide suggests 26B A4B as a flexible general starting point when desktop or server resources are available.
All models accept text and images. E2B, E4B, and 12B also accept audio. Video is analyzed as image frames. All models output text.
E2B and E4B support 128K tokens. The 12B, 26B A4B, and 31B models support 256K tokens.
It can generate structured function calls, but it cannot execute code by itself. The surrounding application must validate arguments, enforce authorization, run the tool, and return results.
There is no per-token license fee for the Apache 2.0 weights. You still pay for hardware or cloud inference, storage, network use, engineering, safety, monitoring, and any third-party service.
Bottom line
Gemma 4 is a strong open-model family for teams that need control over deployment without giving up modern reasoning, multimodal, long-context, coding, and function-calling capabilities. Its biggest advantage is choice: small edge models, a unified multimodal 12B, an efficient 26B A4B, and a larger dense 31B. The right decision comes from benchmarking the smallest viable variant on the actual hardware and application—not from choosing the largest parameter count.
Visit Gemma 4 website ↗
LFM2.5-350M - Liquid AI's 350M-parameter edge model built for tool use and on-device agents

GLM-5.1 — Z AI's new open-source flagship model built for long-horizon agentic coding

Trinity-Large-Thinking - Arcee AI's new open-weight frontier reasoning model for long-horizon agents

Muse Spark - Meta Superintelligence Labs' new multimodal reasoning model

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.