The Rundown AI homepage

Independent tool overview

Gemini 3.1 Flash-Lite at a glance

Gemini 3.1 Flash-Lite is Google's stable, low-cost multimodal model for high-volume translation, transcription, extraction, classification, moderation, and lightweight agent workflows. It accepts text, images, video, audio, and PDFs, returns text or structured output, and supports a 1 million-token context window plus Google tools. It remains available, but Google has announced a May 7, 2027 shutdown and recommends Gemini 3.5 Flash-Lite as the replacement.

Visit the official Gemini 3.1 Flash-Lite site ↗
Gemini 3.1 Flash-Lite product preview
Model ID
gemini-3.1-flash-lite
Lifecycle
Stable, with shutdown scheduled for May 7, 2027
Replacement
Gemini 3.5 Flash-Lite
Inputs
Text, images, video, audio, and PDF
Output
Text, including structured JSON
Token limits
1,048,576 input and 65,536 output
Reviewed
August 31, 2026

Overview

What Gemini 3.1 Flash-Lite is

Gemini 3.1 Flash-Lite prioritizes throughput, latency, and unit economics over maximum frontier reasoning. It fits repetitive workloads where a small per-request saving matters at scale and where failures can be measured, retried, or routed to a stronger model.

The stable endpoint `gemini-3.1-flash-lite` accepts text, images, video, audio, and PDFs, produces text, supports up to 1,048,576 input tokens and 65,536 output tokens, and exposes thinking levels. Its tool set includes caching, code execution, file search, function calling, Search and Maps grounding, structured output, and URL context.

It is not an image, speech, or live-conversation generator. Audio can be supplied for transcription or analysis, but the model returns text; image generation, audio generation, Computer Use, and the Live API are not supported.

The endpoint is stable but already has a retirement date. Google's current deprecation table lists May 7, 2027 as the shutdown date and Gemini 3.5 Flash-Lite as the replacement. New systems should benchmark that successor now, while existing deployments need a scheduled migration rather than waiting for the final notice.

Use cases

Who Gemini 3.1 Flash-Lite is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

High-volume extraction and classification

Parsing reviews, tickets, catalogs, forms, documents, images, and media into validated schemas where unit cost and throughput drive the architecture.

Translation and transcription pipelines

Large batches of short, straightforward language or audio tasks with automated quality checks, terminology constraints, and escalation for uncertain cases.

Moderation assistance

First-pass labeling and prioritization when policy thresholds are explicit and human review remains responsible for ambiguous or consequential enforcement.

Lightweight tool-using agents

Narrow workflows with small allowlisted tools, server-side authorization, schema validation, and a stronger-model or human fallback for difficult decisions.

Capabilities

Core Gemini 3.1 Flash-Lite features

1

Low-cost multimodal input

Processes text, images, video, audio, and PDFs at a lower token price than Google's larger reasoning models.

2

Million-token context

Accepts up to 1,048,576 input tokens and can return as many as 65,536 output tokens, though long inputs still require retrieval and quality testing.

3

Thinking controls

Lets developers adjust reasoning effort to balance speed, cost, and task quality rather than paying the same thinking cost for every request.

4

Structured output and function calling

Can produce schema-constrained JSON and request application functions, making it useful for automated extraction and bounded agent flows.

5

Search, Maps, URL, and file grounding

Supports Google Search and Maps grounding, URL context, and file search for current or supplied evidence, with separate cost and data-retention implications.

6

Code execution

Can use Google's code execution capability for calculations and data-processing steps where supported by the selected API configuration.

7

Caching and multiple consumption modes

Supports explicit context caching plus standard, batch, flex, and priority inference options for different latency, throughput, and reliability requirements.

Process

How the Gemini 3.1 Flash-Lite workflow works

  1. Step 1

    Choose a measurable task

    Use Flash-Lite for bounded work with a clear schema or scoring rubric. Do not begin with an open-ended, high-stakes decision that has no reliable ground truth.

  2. Step 2

    Create a gold evaluation set

    Sample languages, document types, image quality, audio conditions, rare classes, adversarial prompts, missing fields, policy edge cases, and known failures from real traffic.

  3. Step 3

    Constrain the response

    Set a precise system instruction, request structured output, validate every field, cap length, reject unexpected tool calls, and keep authorization outside the model.

  4. Step 4

    Route by confidence and consequence

    Auto-accept only low-risk cases that pass deterministic checks; retry or send ambiguous items to Gemini 3.5 Flash-Lite, a stronger model, or a trained reviewer.

  5. Step 5

    Measure full cost

    Track input, thinking and output tokens, retries, caching, grounding calls, file storage, latency tier, human review, and error remediation rather than comparing input price alone.

  6. Step 6

    Migrate before May 2027

    Benchmark Gemini 3.5 Flash-Lite on the same gold set, test tool and schema behavior, canary traffic, retain rollback, and complete the production switch well before the announced shutdown.

Cost

Gemini 3.1 Flash-Lite pricing and free plan

Standard paid Gemini Developer API pricing is $0.25 per million text, image, or video input tokens, $0.50 per million audio input tokens, and $1.50 per million output tokens including thinking. Free-tier token use is available within quota but Google marks free-tier data as used to improve its products; paid-tier data is marked as not used for that purpose. Caching, grounding, storage, file use, and priority or other consumption modes can add or change cost.

Gemini API free tier

No token charge within quota

For evaluation and limited workloads subject to current rate limits and feature eligibility.

  • Text, visual, audio, and output tokens are free within quota
  • Free-tier data is marked as used to improve Google's products
  • Search and Maps grounding can be tested in AI Studio but are not available as ordinary free-tier production features
  • Quotas and eligible regions can change

Standard paid input

$0.25 per 1M text/image/video tokens

On-demand inference for interactive and general production workloads.

  • Audio input is $0.50 per 1M tokens
  • Output including thinking is $1.50 per 1M tokens
  • Paid-tier data is marked as not used to improve Google's products
  • Actual spend depends on prompt length, reasoning, output, retries, and tools

Explicit context caching

$0.025 per 1M cached text/image/video tokens

Reduces repeated input-processing cost when the same large context is reused.

  • Cached audio is $0.05 per 1M tokens
  • Storage is $1 per 1M tokens per hour
  • Set an appropriate TTL and delete sensitive caches
  • Compare total cache storage plus use against sending fresh context

Grounding tools

5,000 shared queries/month, then $14 per 1,000

Optional paid Google Search or Maps grounding for current and location-aware evidence.

  • Allowance is shared across Gemini 3 models
  • One prompt can trigger multiple billable queries
  • Google Search grounding stores prompts, context, and output for 30 days
  • Retrieved evidence still requires validation

Batch, flex, and priority

Varies by consumption mode

Alternative serving modes trade turnaround, capacity guarantees, and price.

  • Batch fits asynchronous bulk work
  • Flex can lower cost for delay-tolerant traffic
  • Priority costs more for higher service priority
  • Confirm the current table, region, quotas, and service terms before forecasting

Pricing checked . Check current pricing at the source ↗

Assessment

Gemini 3.1 Flash-Lite strengths and limitations

Where it stands out

  • Low standard token prices make it attractive for workloads that run millions of similar requests.
  • Broad multimodal input avoids separate models for many transcription, document, image, and video extraction jobs.
  • Structured output, function calling, and code execution support practical automation beyond plain chat.
  • A 1 million-token window accommodates large source sets and media, while caching can reduce the cost of reused context.
  • Search, Maps, URL, and file grounding provide several ways to supply evidence when the January 2025 knowledge cutoff is insufficient.
  • Stable model naming gives existing users a predictable endpoint during the published migration period.

What to consider

  • Google has scheduled the stable endpoint to shut down on May 7, 2027 and recommends Gemini 3.5 Flash-Lite, so new adoption creates a near-term migration obligation.
  • Flash-Lite is designed for straightforward, high-volume tasks, not the hardest reasoning, coding, planning, or consequential judgments. Low price does not make an inaccurate result inexpensive at scale.
  • The model's January 2025 knowledge cutoff requires current grounding for newer facts, and grounding can still retrieve or interpret weak evidence incorrectly.
  • The 1 million-token limit does not guarantee reliable recall or reasoning across every part of a huge input. Retrieval, chunking, citations, and targeted tests remain necessary.
  • It outputs text only. Image generation, audio generation, live speech, and Computer Use require different models or systems.
  • Generated translations, transcripts, moderation labels, code, tool calls, and structured fields can be wrong. Names, numbers, negation, dialects, policy exceptions, and low-quality media need special testing.
  • Free-tier submitted data may be used to improve Google's products. Paid traffic avoids that training use but can still have limited abuse-monitoring logs and feature-specific retention.
  • Search and Maps grounding add per-query charges and retention; file uploads and explicit caches persist until deletion or expiry and need their own data lifecycle.
  • Automated moderation, employment, education, credit, insurance, healthcare, legal, government, or safety decisions require qualified review, bias testing, appeal paths, and applicable legal authority.

Compare

Gemini 3.1 Flash-Lite alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Gemini 3.5 Flash

Choose Gemini 3.5 Flash when a newer supported model with stronger reasoning and agentic performance justifies a higher cost than the Lite tier.

Explore Gemini 3.5 Flash

Consumer

Gemini 3.7 Flash

Choose Gemini Flash 3.7 for Google's newest GA workhorse when coding, agents, and complex execution matter more than lowest possible unit price.

Explore Gemini 3.7 Flash

Consumer

Gemini 3.1 Pro

Choose Gemini 3.1 Pro for difficult multimodal reasoning or coding that repeatedly fails Flash-Lite evaluation, while accounting for its own preview lifecycle and pricing.

Explore Gemini 3.1 Pro

Questions

Gemini 3.1 Flash-Lite FAQs

What is Gemini 3.1 Flash-Lite best for?

It is best for high-volume, straightforward tasks such as translation, transcription, extraction, classification, moderation assistance, and narrow tool-using workflows where latency and API cost matter.

Is Gemini 3.1 Flash-Lite still active?

Yes. The stable `gemini-3.1-flash-lite` endpoint remains available, but Google has announced a May 7, 2027 shutdown and recommends Gemini 3.5 Flash-Lite as the replacement.

How much does Gemini 3.1 Flash-Lite cost?

Standard paid pricing is $0.25 per million text, image, or video input tokens, $0.50 per million audio input tokens, and $1.50 per million output tokens including thinking. Grounding, caching, storage, and serving modes add separate costs.

Does Gemini 3.1 Flash-Lite support images and audio?

It accepts images, video, audio, PDFs, and text for analysis but returns text. It does not generate images or audio and is not a Live API voice model.

What tools does it support?

The current model page lists caching, code execution, file search, function calling, Google Search and Maps grounding, structured output, thinking, and URL context. It does not support Computer Use, image generation, audio generation, or Live API.

Does Google train on Gemini API data?

Google's pricing page marks free-tier data as used to improve products and paid-tier data as not used for that purpose. Paid requests may still be logged for a limited period for abuse monitoring, and grounding, files, caches, or stored state have separate retention.

Should I start a new project on Gemini 3.1 Flash-Lite?

Only if a benchmark shows a specific short-term advantage and the team accepts migration before May 2027. Otherwise, evaluate Google's recommended 3.5 Flash-Lite successor first.

Can I use it for automated moderation or eligibility decisions?

Use it only as a bounded signal with representative bias and error testing, deterministic policy checks, human review for ambiguous or consequential cases, notice and appeal where required, and no autonomous high-stakes decision without legal authority.

Bottom line

Our Gemini 3.1 Flash-Lite verdict

Gemini 3.1 Flash-Lite is a capable efficiency model with unusually broad input and tool support for its price. It makes sense for measured, reversible pipelines that can validate outputs and route hard cases elsewhere. Its announced May 2027 shutdown changes the recommendation: existing users can keep it while migrating, but new users should benchmark Gemini 3.5 Flash-Lite first and adopt 3.1 only when the near-term cost or compatibility advantage clearly exceeds the migration burden.

Visit Gemini 3.1 Flash-Lite website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.