Maintaining an existing integration
Teams that already depend on `gemini-3-flash-preview` and need accurate current pricing and a controlled migration plan.
Independent tool overview
Gemini 3 Flash was Google's December 2025 speed-focused reasoning model for chat, coding, multimodal analysis, grounding, and agentic applications. The API endpoint remains available as the preview model ID `gemini-3-flash-preview`, but Google now recommends Gemini 3.6 Flash as its replacement and has since released Gemini 3.7 Flash. Treat Gemini 3 Flash as a maintained legacy preview for existing workloads, not the default choice for a new production build.
Visit the official Gemini 3 Flash site ↗
Overview
Gemini 3 Flash launched as a lower-latency, lower-cost way to access much of the Gemini 3 family's reasoning, tool use, coding, and multimodal capability. It became the default model in the Gemini app and AI Mode in Search at launch and was also exposed through Google AI Studio, the Gemini API, Gemini CLI, Antigravity, Vertex AI, and Gemini Enterprise.
The developer endpoint accepts text, image, video, and audio inputs, supports a 1 million-token input context window and up to 64,000 output tokens, and uses dynamic thinking. Developers can constrain the thinking level for faster responses when a task does not justify maximum reasoning cost.
Its current lifecycle matters more than its launch benchmarks. As of August 31, 2026, Google still lists `gemini-3-flash-preview` without a shutdown date, but explicitly recommends `gemini-3.6-flash` as the replacement. Gemini 3.7 Flash is the newer generally available workhorse for coding and agents.
A preview endpoint may change and can receive a short deprecation window. Teams already using Gemini 3 Flash should benchmark a fixed evaluation set against 3.6 or 3.7, remove incompatible sampling parameters, and migrate before Google announces a shutdown.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Teams that already depend on `gemini-3-flash-preview` and need accurate current pricing and a controlled migration plan.
Developers comparing a known Gemini 3 Flash baseline with 3.6 or 3.7 for quality, latency, tool use, token consumption, and regressions.
Noncritical prototypes where the team accepts endpoint changes, restrictive quotas, and eventual migration work.
Capabilities
Modulates reasoning effort by task complexity and supports thinking-level controls, allowing a tradeoff between latency, token use, and answer quality.
Processes up to 1 million input tokens across text, images, video, and audio, with a 64,000-token maximum output.
Was designed for iterative development, tool use, multimodal agents, data extraction, visual question answering, and responsive applications.
Can participate in Gemini API workflows using supported grounding, URL context, code execution, function calling, structured output, and related platform tools; support must be checked for the exact endpoint and API surface.
Originally shipped across the Gemini app, Search AI Mode, AI Studio, Gemini API, Gemini CLI, Antigravity, Vertex AI, and Gemini Enterprise, though consumer defaults have since moved on.
Process
Step 1
Record the exact model ID, SDK and API versions, generation parameters, tools, safety settings, context caching, regions, quotas, and downstream schemas used by the current workload.
Step 2
Include routine, difficult, adversarial, multimodal, long-context, tool-call, refusal, and malformed-input cases with human-approved expected behavior.
Step 3
Run the same set against Gemini 3.6 and 3.7 Flash, measuring task success, latency percentiles, token usage, cost, citation quality, tool reliability, and safety regressions.
Step 4
Google's 3.7 migration guide tells users moving from Gemini 3 Flash to remove deprecated `temperature`, `top_p`, and `top_k` settings, replace `thinking_budget` with `thinking_level`, and remove unsupported prefilled model turns and candidate counts.
Step 5
Pin a specific stable model, release to a small traffic share, preserve rollback, and alert on error rate, malformed tool calls, output-schema failures, cost, latency, refusals, and quality drift.
Cost
Google's current developer guide lists Gemini 3 Flash Preview at $0.50 per million text, image, or video input tokens, $1 per million audio input tokens, and $3 per million output tokens including thinking. Batch pricing on Vertex AI is lower. The Gemini API free tier can be free of token charges but has tighter limits and allows submitted data to improve Google's products; the paid tier says submitted data is not used for that purpose. Grounding, storage, and other tools can add separate charges.
No token charge within free limits
For evaluation and low-volume experimentation subject to available quotas.
$0.50 input / $3 output per 1M tokens
On-demand developer API pricing for text, image, and video input.
$0.50 input / $3 output per 1M tokens globally
Google Cloud access for governed enterprise deployments and regional platform controls.
$0.25 input / $1.50 output per 1M tokens
Discounted asynchronous processing for workloads that do not need interactive latency.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Choose Gemini 3.5 Flash when moving to a newer supported Flash generation while retaining a similar 1M-context, multimodal, thinking, and tool-oriented workflow.
Explore Gemini 3.5 Flash →Consumer
Choose Gemini 3.1 Flash-Lite for high-volume, latency-sensitive workloads where cost matters more than maximum reasoning quality, while noting its announced 2027 shutdown path.
Explore Gemini 3.1 Flash-Lite →Consumer
Choose Gemini 3.1 Pro for harder reasoning and multimodal work where quality is more important than Flash latency and cost.
Explore Gemini 3.1 Pro →Questions
Yes. Google still listed `gemini-3-flash-preview` without an announced shutdown date on August 31, 2026. It is a preview endpoint and Google recommends Gemini 3.6 Flash as its replacement.
Usually no. Benchmark Gemini 3.6 and the newer GA Gemini 3.7 Flash first. Starting on a superseded preview creates unnecessary migration risk unless a measured compatibility or cost requirement justifies it.
Current paid API pricing is $0.50 per million text, image, or video input tokens, $1 per million audio input tokens, and $3 per million output tokens including thinking. Batch processing is cheaper, and tools can add charges.
Google lists a 1 million-token input context window and a 64,000-token maximum output. Practical quality across very long context still needs workload-specific testing.
Google's pricing table marks free-tier submissions as used to improve its products and paid-tier submissions as not used for that purpose. Verify the applicable API or Cloud terms rather than assuming the Gemini consumer app's privacy settings apply.
Google's deprecation table names Gemini 3.6 Flash as the recommended replacement. Gemini 3.7 Flash is the newer generally available workhorse announced in August 2026, so compare both against the existing workload.
Model behavior, token use, tool calls, safety responses, structured output, and latency can all shift. Google also documents request changes for 3.7, including removal of deprecated sampling parameters and prefilled model turns and use of `thinking_level` instead of `thinking_budget`.
Bottom line
Gemini 3 Flash was an important speed-and-reasoning release and remains usable for existing integrations, but its current value is as a migration baseline. It is still a preview, Google already names 3.6 Flash as the replacement, and 3.7 Flash is now GA. Keep it only while a measured test shows a real advantage, protect the workload with a pinned endpoint and rollback, and schedule migration before a shutdown date turns routine maintenance into an outage.
Visit Gemini 3 Flash website ↗
CC - Google Labs’ experimental AI productivity agent in Gmail

Chatterbox Turbo - Resemble AI's fast, expressive, open-source text-to-speech model

Disco - Google's experimental browser that uses Gemini 3 to generate custom web applications based on your open tabs and browsing tasks

MiMo-V2-Flash - Xiaomi's powerful open-weights reasoning model

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.