Teams needing local translation
Downloaded weights can keep inference inside a controlled environment when the deployment is secured and properly operated.
Independent tool overview
TranslateGemma is Google's family of downloadable translation models based on Gemma 3, offered in 4B, 12B, and 27B sizes for text and text-in-image translation across 55 evaluated languages.
Visit the official TranslateGemma site ↗
Overview
TranslateGemma is a model family for developers and researchers, not a polished translation website for consumers. Google publishes 4B, 12B, and 27B instruction-tuned weights that accept source and target language codes plus either text or an image, then generate translated text. The models can be downloaded from Hugging Face or Kaggle or deployed through compatible local, cloud, and Google Cloud tooling.
The core release covers 55 languages evaluated on WMT24++ and retains Gemma 3's vision input for translating text found in images. Google also trained on nearly 500 additional language pairs, but explicitly says it did not yet have confirmed evaluation metrics for that extended set at launch. The model card documents a 2,000-token total input context, so this is not a drop-in system for translating arbitrarily long documents.
TranslateGemma's efficiency results are promising, but benchmark averages do not certify any particular language pair, domain, dialect, or document. Production teams still need bilingual evaluation, terminology controls, privacy architecture, fallback behavior, and human review—especially for legal, medical, financial, safety, immigration, or other consequential content.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Downloaded weights can keep inference inside a controlled environment when the deployment is secured and properly operated.
Three sizes, published evaluation data, and a technical report make the family useful for benchmarking, adaptation, and language-pair research.
Google positions the 4B model for mobile or edge work and the 12B model for consumer laptops, subject to runtime, precision, and hardware constraints.
The multimodal interface can extract and translate visible text from an image for constrained use cases.
The models can be evaluated or adapted for a defined language pair and terminology set under the applicable Gemma terms.
Capabilities
Provides 4B, 12B, and 27B variants so developers can trade memory, throughput, latency, and translation quality.
Google reports WMT24++ results across 55 languages spanning high-, medium-, and low-resource families.
Accepts a text string with explicit source and target language codes and returns translated text.
Accepts an image and returns translated text using the retained Gemma 3 multimodal capability.
The official chat template supports two-letter language codes and regional variants such as en-GB.
Official checkpoints are distributed through Hugging Face and Kaggle after acceptance of the Gemma usage terms.
The official model page documents Transformers and compatible serving options, while Google offers deployment through Vertex AI Model Garden.
The model card reports MetricX, COMET, MQM, and image-translation results for specified datasets and variants.
Process
Step 1
Define source and target locales, content domain, expected volume, and the consequence of an incorrect translation.
Step 2
Read the current Gemma Terms and Prohibited Use Policy, including redistribution requirements, before downloading or serving the model.
Step 3
Benchmark 4B, 12B, and 27B on target hardware rather than assuming Google's broad deployment labels guarantee acceptable latency or quality.
Step 4
Pass exactly one text or image item with supported source and target language codes through the official chat template.
Step 5
Build a bilingual test set containing terminology, names, numbers, formatting, ambiguity, dialect, sensitive content, and adversarial inputs from the actual use case.
Step 6
Implement chunking, glossary or translation-memory logic, confidence or review routing, observability, access control, and safe failure behavior.
Step 7
Have qualified reviewers approve consequential translations and continuously sample live output for regressions across each supported locale.
Cost
Google distributes TranslateGemma model weights under the Gemma Terms rather than selling a TranslateGemma subscription. Access through Hugging Face requires accepting the license. The real cost is hardware, cloud inference, storage, engineering, monitoring, and human quality review; managed Vertex AI charges depend on the infrastructure and deployment configuration.
Weights available under Gemma terms
Smallest variant, positioned by Google for mobile and edge deployment.
Weights available under Gemma terms
Middle variant, positioned for consumer laptops and local development.
Weights available under Gemma terms
Largest official variant for maximum fidelity in the release family.
Usage and infrastructure charges vary
Serve through Vertex AI Model Garden or another compatible cloud stack.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Choose Tiny Aya when a compact multilingual generative model with broader advertised language coverage is more important than translation-specific tuning.
Explore Tiny Aya →Consumer
Choose Translate with ChatGPT when end users need a ready-made text, voice, and image translation interface rather than self-hosted model weights.
Explore Translate with ChatGPT →Miscellaneous
Choose Gemini Live Translate when the primary requirement is real-time spoken translation instead of batch text or image inference.
Explore Gemini 3.5 Live Translate →Questions
TranslateGemma is a family of Google translation models based on Gemma 3. Developers can download 4B, 12B, and 27B variants to translate text or visible text in images.
Not directly. Google Translate is a finished consumer service. TranslateGemma is a set of model weights and documentation for developers who will build, host, evaluate, and operate their own translation workflow.
Google reports core training and evaluation across 55 languages. It also trained on nearly 500 additional language pairs but said confirmed evaluation metrics were not yet available for that extended set at launch.
Start with 4B for constrained hardware, 12B for a balance of quality and cost, and 27B for the highest reported fidelity. Benchmark every candidate on your own language pair, domain, hardware, latency target, and review process.
Yes. Google positions 4B for mobile or edge use and 12B for consumer laptops. Actual feasibility depends on runtime, precision or quantization, available memory, accelerator support, and performance requirements.
Yes. The models accept an image and return translated text. The published image benchmark used a constrained set with a single text region, so complex documents and layouts require separate testing.
The official weights are downloadable after accepting the Gemma terms; there is no separate TranslateGemma subscription listed. Running the model still incurs hardware, cloud, engineering, monitoring, and human-review costs.
Google calls it an open model and publishes weights, but use and redistribution are governed by the Gemma Terms and Prohibited Use Policy. Teams should review those terms rather than assuming a standard unrestricted open-source software license.
The official model card states a total input context of 2,000 tokens. Longer documents need structure-aware segmentation and a process for preserving terminology, references, formatting, and cross-segment context.
Do not rely on benchmark averages for consequential material. Use a qualified professional translator or reviewer, preserve the source, test the exact language direction and domain, and prohibit automatic release when an error could affect rights, care, money, or safety.
Local inference can keep source content off an external model API, but privacy still depends on the complete system: logs, telemetry, storage, backups, access controls, updates, and any surrounding services.
Use native bilingual reviewers and representative material covering terminology, ambiguity, names, numbers, negation, units, dialect, formatting, offensive content, injection attempts, and known failure cases. Measure each language direction separately.
Bottom line
TranslateGemma is a compelling building block for teams that need downloadable, translation-focused models and are willing to operate the full quality and serving stack. The 12B model's efficiency and the availability of 4B and 27B variants make experimentation practical. The responsible path is narrow and empirical: choose one language direction, benchmark real content, add terminology and review controls, and expand only after native speakers verify performance.
Visit TranslateGemma website ↗
Slackbot - Slack's personal AI agent for work

Step3-VL-10B - StepFun's open-source SOTA vision language model

Gmail - Google's email inbox, now infused with Gemini for AI-powered insights, actions, and improvements

ERNIE 5.0 - Baidu's omni-modal, top-ranked Chinese model

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.