The Rundown AI homepage

Independent tool overview

Gemini 3 Flash at a glance

Gemini 3 Flash was Google's December 2025 speed-focused reasoning model for chat, coding, multimodal analysis, grounding, and agentic applications. The API endpoint remains available as the preview model ID `gemini-3-flash-preview`, but Google now recommends Gemini 3.6 Flash as its replacement and has since released Gemini 3.7 Flash. Treat Gemini 3 Flash as a maintained legacy preview for existing workloads, not the default choice for a new production build.

Visit the official Gemini 3 Flash site ↗
Gemini 3 Flash product preview
Model ID
gemini-3-flash-preview
Lifecycle
Preview; available, but superseded for new builds
Recommended replacement
Gemini 3.6 Flash; Gemini 3.7 Flash is the newer GA workhorse
Input context
1 million tokens
Maximum output
64,000 tokens
Knowledge cutoff
January 2025 in Google's current model table
Reviewed
August 31, 2026

Overview

What Gemini 3 Flash is

Gemini 3 Flash launched as a lower-latency, lower-cost way to access much of the Gemini 3 family's reasoning, tool use, coding, and multimodal capability. It became the default model in the Gemini app and AI Mode in Search at launch and was also exposed through Google AI Studio, the Gemini API, Gemini CLI, Antigravity, Vertex AI, and Gemini Enterprise.

The developer endpoint accepts text, image, video, and audio inputs, supports a 1 million-token input context window and up to 64,000 output tokens, and uses dynamic thinking. Developers can constrain the thinking level for faster responses when a task does not justify maximum reasoning cost.

Its current lifecycle matters more than its launch benchmarks. As of August 31, 2026, Google still lists `gemini-3-flash-preview` without a shutdown date, but explicitly recommends `gemini-3.6-flash` as the replacement. Gemini 3.7 Flash is the newer generally available workhorse for coding and agents.

A preview endpoint may change and can receive a short deprecation window. Teams already using Gemini 3 Flash should benchmark a fixed evaluation set against 3.6 or 3.7, remove incompatible sampling parameters, and migrate before Google announces a shutdown.

Use cases

Who Gemini 3 Flash is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Maintaining an existing integration

Teams that already depend on `gemini-3-flash-preview` and need accurate current pricing and a controlled migration plan.

Migration benchmarking

Developers comparing a known Gemini 3 Flash baseline with 3.6 or 3.7 for quality, latency, tool use, token consumption, and regressions.

Temporary preview experiments

Noncritical prototypes where the team accepts endpoint changes, restrictive quotas, and eventual migration work.

Capabilities

Core Gemini 3 Flash features

1

Dynamic thinking

Modulates reasoning effort by task complexity and supports thinking-level controls, allowing a tradeoff between latency, token use, and answer quality.

2

Long multimodal context

Processes up to 1 million input tokens across text, images, video, and audio, with a 64,000-token maximum output.

3

Coding and agentic reasoning

Was designed for iterative development, tool use, multimodal agents, data extraction, visual question answering, and responsive applications.

4

Grounding and tools

Can participate in Gemini API workflows using supported grounding, URL context, code execution, function calling, structured output, and related platform tools; support must be checked for the exact endpoint and API surface.

5

Multiple Google surfaces

Originally shipped across the Gemini app, Search AI Mode, AI Studio, Gemini API, Gemini CLI, Antigravity, Vertex AI, and Gemini Enterprise, though consumer defaults have since moved on.

Process

How the Gemini 3 Flash workflow works

  1. Step 1

    Inventory dependencies

    Record the exact model ID, SDK and API versions, generation parameters, tools, safety settings, context caching, regions, quotas, and downstream schemas used by the current workload.

  2. Step 2

    Build a representative evaluation set

    Include routine, difficult, adversarial, multimodal, long-context, tool-call, refusal, and malformed-input cases with human-approved expected behavior.

  3. Step 3

    Compare newer stable targets

    Run the same set against Gemini 3.6 and 3.7 Flash, measuring task success, latency percentiles, token usage, cost, citation quality, tool reliability, and safety regressions.

  4. Step 4

    Update incompatible requests

    Google's 3.7 migration guide tells users moving from Gemini 3 Flash to remove deprecated `temperature`, `top_p`, and `top_k` settings, replace `thinking_budget` with `thinking_level`, and remove unsupported prefilled model turns and candidate counts.

  5. Step 5

    Canary and monitor

    Pin a specific stable model, release to a small traffic share, preserve rollback, and alert on error rate, malformed tool calls, output-schema failures, cost, latency, refusals, and quality drift.

Cost

Gemini 3 Flash pricing and free plan

Google's current developer guide lists Gemini 3 Flash Preview at $0.50 per million text, image, or video input tokens, $1 per million audio input tokens, and $3 per million output tokens including thinking. Batch pricing on Vertex AI is lower. The Gemini API free tier can be free of token charges but has tighter limits and allows submitted data to improve Google's products; the paid tier says submitted data is not used for that purpose. Grounding, storage, and other tools can add separate charges.

Gemini API free tier

No token charge within free limits

For evaluation and low-volume experimentation subject to available quotas.

  • Free limits and availability can change
  • Preview models can have more restrictive quotas
  • Google's pricing table marks free-tier submitted data as used to improve products
  • Not a substitute for enterprise data terms
  • Grounding and some tools may be unavailable or separately limited

Gemini API paid standard

$0.50 input / $3 output per 1M tokens

On-demand developer API pricing for text, image, and video input.

  • Audio input is $1 per 1M tokens
  • Output price includes thinking tokens
  • Paid-tier submitted data is marked as not used to improve products
  • Search grounding and other tools can add charges
  • Taxes, currency, region, quotas, and platform terms can differ

Vertex AI standard

$0.50 input / $3 output per 1M tokens globally

Google Cloud access for governed enterprise deployments and regional platform controls.

  • Audio input is $1 per 1M tokens
  • Provisioned throughput and non-global terms can differ
  • Cloud logging, storage, grounding, networking, and tools may add cost
  • Review the Cloud agreement and data residency configuration

Vertex AI batch

$0.25 input / $1.50 output per 1M tokens

Discounted asynchronous processing for workloads that do not need interactive latency.

  • Audio input is $0.50 per 1M tokens
  • Batch availability and turnaround are separate from online serving
  • Not suitable for immediate user-facing responses
  • Validate current regional availability and quotas

Pricing checked . Check current pricing at the source ↗

Assessment

Gemini 3 Flash strengths and limitations

Where it stands out

  • The 1 million-token context and multimodal inputs made it useful for large documents, codebases, recordings, images, and video in one request.
  • Dynamic thinking lets simple requests complete quickly while harder requests can consume more reasoning effort.
  • Launch pricing was inexpensive relative to larger frontier models, especially for text and visual input.
  • The endpoint supports the broader Gemini tool and grounding ecosystem used by many agentic applications.
  • It remains available, giving existing users time to test and migrate rather than forcing an immediate emergency change.

What to consider

  • Gemini 3 Flash is a preview model, not a fixed generally available production target. Google recommends Gemini 3.6 Flash as its replacement.
  • The newer Gemini 3.7 Flash is GA and specifically positioned as Google's more capable workhorse for coding and agents, so choosing 3 Flash for a new system creates avoidable migration debt.
  • Preview endpoints can change behavior, quotas, pricing, tool support, or availability and may receive only a short deprecation window.
  • The January 2025 knowledge cutoff makes grounding or an authoritative external data source necessary for current facts, and grounding can still retrieve weak or misread evidence.
  • Thinking tokens are billed as output, so a low headline input price does not predict total request cost for complex prompts.
  • Model-generated code, tool calls, citations, and structured output can be wrong or unsafe; schemas, permissions, tests, and human approval remain necessary.
  • Free-tier Gemini API data may be used to improve Google's products. Confidential or regulated workloads need appropriate paid or enterprise terms and a documented security review.
  • Google's launch benchmark figures are vendor-reported snapshots, not guarantees for a specific workload, language, region, latency target, or production tool chain.

Compare

Gemini 3 Flash alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Gemini 3.5 Flash

Choose Gemini 3.5 Flash when moving to a newer supported Flash generation while retaining a similar 1M-context, multimodal, thinking, and tool-oriented workflow.

Explore Gemini 3.5 Flash

Consumer

Gemini 3.1 Flash-Lite

Choose Gemini 3.1 Flash-Lite for high-volume, latency-sensitive workloads where cost matters more than maximum reasoning quality, while noting its announced 2027 shutdown path.

Explore Gemini 3.1 Flash-Lite

Consumer

Gemini 3.1 Pro

Choose Gemini 3.1 Pro for harder reasoning and multimodal work where quality is more important than Flash latency and cost.

Explore Gemini 3.1 Pro

Questions

Gemini 3 Flash FAQs

Is Gemini 3 Flash still available?

Yes. Google still listed `gemini-3-flash-preview` without an announced shutdown date on August 31, 2026. It is a preview endpoint and Google recommends Gemini 3.6 Flash as its replacement.

Should I start a new project on Gemini 3 Flash?

Usually no. Benchmark Gemini 3.6 and the newer GA Gemini 3.7 Flash first. Starting on a superseded preview creates unnecessary migration risk unless a measured compatibility or cost requirement justifies it.

How much does Gemini 3 Flash cost?

Current paid API pricing is $0.50 per million text, image, or video input tokens, $1 per million audio input tokens, and $3 per million output tokens including thinking. Batch processing is cheaper, and tools can add charges.

What is Gemini 3 Flash's context window?

Google lists a 1 million-token input context window and a 64,000-token maximum output. Practical quality across very long context still needs workload-specific testing.

What data does the Gemini API use for training?

Google's pricing table marks free-tier submissions as used to improve its products and paid-tier submissions as not used for that purpose. Verify the applicable API or Cloud terms rather than assuming the Gemini consumer app's privacy settings apply.

What replaces Gemini 3 Flash?

Google's deprecation table names Gemini 3.6 Flash as the recommended replacement. Gemini 3.7 Flash is the newer generally available workhorse announced in August 2026, so compare both against the existing workload.

What can break during migration?

Model behavior, token use, tool calls, safety responses, structured output, and latency can all shift. Google also documents request changes for 3.7, including removal of deprecated sampling parameters and prefilled model turns and use of `thinking_level` instead of `thinking_budget`.

Bottom line

Our Gemini 3 Flash verdict

Gemini 3 Flash was an important speed-and-reasoning release and remains usable for existing integrations, but its current value is as a migration baseline. It is still a preview, Google already names 3.6 Flash as the replacement, and 3.7 Flash is now GA. Keep it only while a measured test shows a real advantage, protect the workload with a pinned endpoint and rollback, and schedule migration before a shutdown date turns routine maintenance into an outage.

Visit Gemini 3 Flash website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.