High-throughput coding agents
Agent loops that need long context, function calls, code execution and lower latency than a large flagship model.
Independent tool overview
Gemini 3.5 Flash is Google's stable multimodal, text-output model for fast coding and agentic workloads, with a one-million-token input window, tool use and paid API rates of $1.50 input and $9 output per million tokens.
Visit the official Gemini 3.5 Flash site ↗
Overview
Gemini 3.5 Flash remains listed as a stable Gemini API model and is available through Google AI Studio and Google's enterprise and developer surfaces. It accepts text, images, video, audio and PDFs, then returns text. The model is designed for coding, multimodal understanding and multi-step agent loops rather than image, audio or real-time voice generation.
There is now a newer migration target. Google's current documentation recommends Gemini 3.7 Flash for workloads moving from 3.5 Flash and says the upgrade improves coding, spatial and multimodal reasoning, design adherence and agent reliability. Existing 3.5 integrations do not need to be labeled inactive, but new projects should benchmark 3.7 before locking 3.5 into production.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Agent loops that need long context, function calls, code execution and lower latency than a large flagship model.
Applications reasoning across text, screenshots, video, audio and PDFs while producing structured text.
Paid applications that need Google Search, Maps, URL context or file search integrated with a reasoning model.
Teams maintaining a stable model ID while they test output quality, latency and compatibility against Gemini 3.7 Flash.
Capabilities
Accepts up to 1,048,576 input tokens and can return as many as 65,536 output tokens.
Processes text, images, video, audio and PDF documents in one model while returning text.
Can spend internal reasoning tokens on harder tasks; those tokens are included in the paid output-token charge.
Supports tool schemas and machine-readable responses for application and agent orchestration.
Supports code execution and offers computer use in preview for workflows that need to act rather than only answer.
Paid API workflows can connect responses to Google Search and Maps, with separate usage charges after the included allowance.
Can retrieve from indexed files and supplied URLs as part of a longer reasoning or agent workflow.
Offers discounted asynchronous or flexible capacity and a higher-priced priority service in addition to standard inference.
Process
Step 1
For a new system, benchmark Gemini 3.7 Flash first because Google now names it as the upgrade target from 3.5.
Step 2
Include real prompts, difficult edge cases, tool failures, long-context examples and expected structured outputs.
Step 3
Cap output, agent steps, searches and retries because thinking tokens and repeated tool calls affect both latency and cost.
Step 4
Use Search, Maps, file retrieval or URL context for claims that require current or private evidence, then preserve source attribution.
Step 5
Check schemas, permissions, code execution and computer-use outcomes before allowing the model to affect shared or production systems.
Step 6
Remove deprecated sampling parameters and prefilled model turns, compare behavior and cost, then roll forward behind monitoring and rollback.
Cost
Gemini 3.5 Flash has a free developer tier and four paid consumption modes. Standard costs $1.50 per million input tokens and $9 per million output tokens including thinking; Batch and Flex halve inference rates, while Priority charges 1.8 times Standard.
$0 within limits
Limited developer access and free Google AI Studio testing.
$1.50 input / $9 output
On-demand paid inference per one million tokens.
$0.75 input / $4.50 output
Discounted asynchronous processing per one million tokens.
$0.75 input / $4.50 output
Discounted flexible-capacity inference per one million tokens.
$2.70 input / $16.20 output
Higher-priced capacity for workloads that need priority service.
5,000 free, then $14/1,000 queries
Separate paid-tool allowance shared across Gemini 3.x models.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Choose Claude Sonnet 4.6 to compare long-context coding, agent behavior and a different safety and platform stack.
Explore Claude Sonnet 4.6 →Consumer
Choose Grok 4.6 for another agentic model with a large context window and integrated SpaceXAI tool ecosystem.
Explore Grok 4.6 →Business Operations
Choose ChatGPT when the priority is a broad end-user workspace rather than direct model integration.
Explore ChatGPT →Questions
Yes. Google currently lists gemini-3.5-flash as a stable model with active free and paid pricing.
It has not been removed, but Gemini 3.7 Flash is Google's recommended migration target for Gemini 3.5 Flash workloads and is the better starting benchmark for new projects.
Standard paid pricing is $1.50 per million input tokens and $9 per million output tokens, including thinking tokens. Batch and Flex are $0.75 input and $4.50 output; Priority is $2.70 input and $16.20 output.
Yes, within published rate limits. Google says free-tier content may be used to improve its products, while paid-tier content is not.
The model supports up to 1,048,576 input tokens and 65,536 output tokens.
It accepts text, images, video, audio and PDFs and returns text. It does not generate images or audio and does not support the Live API.
Yes. It supports function calling, structured outputs, code execution, file search, URL context, Google Search and Maps grounding, plus computer use in preview.
Paid Gemini 3.x models share 5,000 free search requests per month, then charge $14 per 1,000 queries. A single model request can produce multiple billable search queries.
Google says migrations from 3.5 Flash should remove deprecated temperature, top-p and top-k sampling parameters and prefilled model turns, then validate the changed behavior.
Start by evaluating Gemini 3.7 Flash because Google recommends it as the migration target. Keep 3.5 only when its current behavior, availability or an incomplete migration gives it a measured advantage.
Bottom line
Gemini 3.5 Flash is still a capable, stable model for multimodal and agentic systems, but it now sits in a transition window. Existing teams can keep it while they run a proper regression and cost test; new teams should begin with Gemini 3.7 Flash. The biggest implementation risks are hidden thinking and search costs, preview computer use, free-tier data handling and assuming a huge context window removes the need for retrieval and evaluation.
Visit Gemini 3.5 Flash website ↗
Alexa for Shopping - Amazon's new shopping agent for Q&A, price tracking, and Auto-Buy across devices

Qwen 3.7 Max- Alibaba's new flagship model for long-horizon agentic tasks

Incognito Chat - Meta’s new way to have private conversations with Meta AI via WhatsApp

Claude Opus 4.8 - Anthropic's new top model with improvements to reliability, coding, and agentic flows.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.