Gemini 3.6 Flash and Gemini 3.5 Flash-Lite at a Glance
Google released both models on July 21, 2026. Unlike many recent Gemini launches, neither model entered the API as a temporary preview. Google lists both as stable and generally available for production use.
| Specification |
Gemini 3.6 Flash |
Gemini 3.5 Flash-Lite |
| Release date |
July 21, 2026 |
July 21, 2026 |
| Availability |
GA / Stable |
GA / Stable |
| Model ID |
gemini-3.6-flash |
gemini-3.5-flash-lite |
| Primary role |
General-purpose workhorse |
High-throughput execution model |
| Default thinking level |
Medium |
Minimal |
| Input context limit |
1,048,576 tokens |
1,048,576 tokens |
| Maximum output |
65,536 tokens |
65,536 tokens |
| Input types |
Text, image, video, audio, PDF |
Text, image, video, audio, PDF |
| Output type |
Text |
Text |
| Standard input price |
$1.50 per 1M tokens |
$0.30 per 1M tokens |
| Standard output price |
$7.50 per 1M tokens |
$2.50 per 1M tokens |
| Best suited for |
Coding, complex agents, knowledge work, spatial and multimodal reasoning |
Extraction, classification, translation, search, data processing, subagents |
Both models have a March 2026 knowledge cutoff, according to their official model cards. For more recent information, developers need to use search grounding or provide current data in the prompt.
What Is Gemini 3.6 Flash?
Gemini 3.6 Flash is Google’s latest general-purpose Flash model for coding, knowledge work, multimodal analysis, and multi-step agentic execution.
It is based on Gemini 3.5 Flash rather than being an entirely separate generation of the Gemini architecture. The main goal of the update is to make Flash more efficient during real work: fewer unnecessary output tokens, fewer reasoning steps, fewer tool calls, and fewer repeated execution loops.
That distinction matters. Gemini 3.6 Flash is not simply trying to produce a higher score on every intelligence benchmark. It is trying to complete the same or better work with less computation and less waiting.
According to Google’s launch announcement, Gemini 3.6 Flash uses approximately 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index. Google also reports that it takes fewer turns and tool calls to complete multi-step workflows.
Official evaluations show improvements in several practical areas:
- DeepSWE increased from 37% to 49%, suggesting fewer incorrect code edits and execution loops.
- MLE-Bench increased from 49.7% to 63.9% for machine learning research tasks.
- OSWorld-Verified increased from 78.4% to 83% for computer-use tasks.
- GDPval-AA v2 increased from 1349 to 1421 for knowledge-work performance.
These are vendor-reported results, so they should not be treated as universal guarantees. However, they support Google’s positioning of Gemini 3.6 Flash as a more reliable model for code migrations, technical diagnostics, document analysis, chart interpretation, and tool-using agents.
The model accepts text, images, video, audio, and PDF files, but it only generates text. It does not directly generate images, videos, or speech.
What Is Gemini 3.5 Flash-Lite—and What Does GA Mean?
Gemini 3.5 Flash-Lite is Google’s fastest and least expensive model in the Gemini 3.5 family. It is optimized for low-latency, high-volume tasks where throughput and API cost matter more than maximum reasoning quality.
The “GA” in Flash-Lite GA means Generally Available. It is an availability status, not part of the model’s name. A GA model is intended for stable production use rather than short-term preview testing.
Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite—not Gemini 3.6 Flash. Therefore, it should not be understood as a compressed version of the new 3.6 model. The two releases come from different upgrade paths:
- Gemini 3.5 Flash → Gemini 3.6 Flash
- Gemini 3.1 Flash-Lite → Gemini 3.5 Flash-Lite
Google positions Flash-Lite for use cases such as:
- Document classification and structured extraction
- Receipt and invoice processing
- Product attribute extraction
- Translation and localization
- Search-result processing
- High-volume customer-support routing
- Tabular data processing
- Repetitive tool calls
- Parallel subagent execution
Independent testing cited in Google’s announcement measured Gemini 3.5 Flash-Lite at approximately 350 output tokens per second. Actual API performance will vary with prompt length, thinking level, server load, tools, and region, but the result illustrates the model’s throughput-focused design.
Flash-Lite is also considerably more capable than its predecessor. Google reports improvements from 31% to 54% on Terminal-Bench 2.1 and from 60.1% to 72.2% on its long-context evaluation.
It can use higher thinking levels for more complicated tasks, but doing so changes its cost and latency profile. If every request requires deep reasoning, Flash-Lite may lose some of the economic advantage that makes it attractive.
What Do “Flash” and “Flash-Lite” Mean in Gemini?
“Flash” is not an official acronym. It is Google’s product label for Gemini models that prioritize speed, efficiency, and scalable inference.
The practical meaning of Flash is:
- Faster than heavier flagship models
- Less expensive to run at scale
- Suitable for interactive applications
- Still capable of reasoning, coding, and multimodal understanding
Flash-Lite moves further toward the efficiency end of that spectrum. It is designed for workloads with a large number of relatively bounded tasks, such as extracting fields from thousands of documents or translating millions of short content items.
“Lite” does not mean that the model cannot reason or process multimodal inputs. Gemini 3.5 Flash-Lite still supports thinking, function calling, structured output, search grounding, and a 1M-token context window.
The difference is how Google expects each model to be deployed:
- Flash: Choose it when the model must understand a complicated goal and decide what to do.
- Flash-Lite: Choose it when the task is already defined and needs to be completed quickly at scale.
Gemini 3.6 Flash vs Gemini 3.5 Flash: What Actually Improved?
Gemini 3.6 Flash is the direct successor to Gemini 3.5 Flash, but “successor” does not mean that every intelligence score has increased.
The most meaningful improvements are efficiency, coding reliability, tool use, and total time per task.
| Comparison |
Gemini 3.6 Flash |
Gemini 3.5 Flash |
| Standard input price |
$1.50/M |
$1.50/M |
| Standard output price |
$7.50/M |
$9/M |
| Default thinking level |
Medium |
Medium |
| Context window |
1M |
1M |
| Maximum output |
64K |
64K |
| Artificial Analysis Intelligence Index |
50 |
50 |
| Average time per task in launch testing |
1.3 minutes |
2.7 minutes |
| Primary advantage |
Lower token use and faster task completion |
Established production baseline |
In Artificial Analysis testing, both models scored 50 on its Intelligence Index. That makes it difficult to argue that Gemini 3.6 Flash represents a large increase in general intelligence.
However, Gemini 3.6 Flash completed the tested tasks in less than half the average time. Its output price is also approximately 16.7% lower, while the model uses fewer output tokens in many agentic workloads.
For production users, this may be more valuable than a small benchmark increase. A model that reaches a similar answer with fewer tokens, fewer failed tool calls, and fewer correction loops can reduce both the infrastructure cost and the time users spend waiting.
There is little pricing incentive to begin a new deployment on Gemini 3.5 Flash when Gemini 3.6 Flash has the same standard input price and a lower output price. Existing applications, however, should still test migration compatibility instead of changing the model ID without review.
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Pricing
The following prices are Google’s paid Gemini API rates per one million tokens at launch.
| Model |
Standard Input |
Standard Output |
Batch Input |
Batch Output |
| Gemini 3.6 Flash |
$1.50 |
$7.50 |
$0.75 |
$3.75 |
| Gemini 3.5 Flash |
$1.50 |
$9.00 |
$0.75 |
$4.50 |
| Gemini 3.5 Flash-Lite |
$0.30 |
$2.50 |
$0.15 |
$1.25 |
| Gemini 3.1 Flash-Lite* |
$0.25 |
$1.50 |
$0.125 |
$0.75 |
*Gemini 3.1 Flash-Lite charges a higher input rate for audio. Refer to the current Gemini API pricing page before deploying a production workload.
A simple cost example
Suppose a workload consumes 100 million input tokens and generates 20 million output tokens.
Ignoring caching, grounding, and storage fees, the standard API cost would be:
- Gemini 3.6 Flash: $150 input + $150 output = $300
- Gemini 3.5 Flash: $150 input + $180 output = $330
- Gemini 3.5 Flash-Lite: $30 input + $50 output = $80
This illustrates the large sticker-price difference between Flash and Flash-Lite. However, fixed token calculations do not tell the entire story.
A cheaper model can become expensive if it requires more retries, produces longer reasoning traces, or fails to complete a task. Similarly, a model with a higher token rate may have a lower cost per completed task if it uses fewer turns and tools.
That is why Gemini 3.6 Flash’s token efficiency matters. If it produces approximately 17% fewer output tokens while also charging 16.7% less for each output token, the savings on the output portion of some workloads can be substantially greater than the price-table difference alone suggests.
Gemini 3.5 Flash-Lite presents the opposite lesson. It is the least expensive model in the 3.5 family, but it is not cheaper than 3.1 Flash-Lite for every token type. Its text, image, and video input price increased from $0.25 to $0.30 per million tokens, while output increased from $1.50 to $2.50.
Developers are paying more than they did for 3.1 Flash-Lite, but they are receiving much better reasoning, agentic performance, and throughput.
Which Gemini Flash Model Should You Choose?
The best model depends on whether your bottleneck is task complexity or execution volume.
| Workload |
Recommended model |
Why |
| Large codebase changes |
Gemini 3.6 Flash |
Better planning and fewer unwanted edits |
| Multi-step coding agent |
Gemini 3.6 Flash |
Stronger tool use and execution loops |
| Chart and document analysis |
Gemini 3.6 Flash |
Better multimodal and spatial reasoning |
| Complex business research |
Gemini 3.6 Flash |
Stronger knowledge-work performance |
| Document classification |
Gemini 3.5 Flash-Lite |
Lower cost and higher throughput |
| Receipt or invoice extraction |
Gemini 3.5 Flash-Lite |
Designed for structured document processing |
| Large-scale translation |
Gemini 3.5 Flash-Lite |
Fast, multimodal, and inexpensive |
| Search-result processing |
Gemini 3.5 Flash-Lite |
Suitable for parallel high-volume execution |
| Simple subagent tasks |
Gemini 3.5 Flash-Lite |
Low-cost execution with adjustable thinking |
| Agent orchestration |
Both |
Flash plans; Flash-Lite executes |
The more interesting answer: use both
Imagine an e-commerce platform that needs to extract product information from thousands of listings, PDFs, images, and supplier documents.
Gemini 3.6 Flash could:
- Inspect several representative documents.
- Design the extraction schema.
- Decide which sources are reliable.
- Identify ambiguous or conflicting product information.
- Review exceptions that require deeper reasoning.
Gemini 3.5 Flash-Lite could then run in parallel to:
- Read every product document.
- Extract brand, material, size, price, and availability.
- Translate product descriptions.
- Return structured JSON.
- Flag unusual records for the main agent.
This architecture avoids paying 3.6 Flash rates for every repetitive extraction while still using the stronger model where its reasoning matters.
Gemini 3.6 Flash vs Kimi K3
Gemini 3.6 Flash and Kimi K3 were released within days of each other, but they optimize for different things.
Kimi K3 is Moonshot AI’s 2.8-trillion-parameter flagship model for long-horizon coding, reasoning, and end-to-end knowledge work. Gemini 3.6 Flash emphasizes speed, token efficiency, multimodal workflows, and integration with Google’s tools.
The following figures come from the same Artificial Analysis comparison, making them more useful than comparing unrelated vendor benchmark tables.
| Comparison |
Gemini 3.6 Flash |
Kimi K3 |
| Intelligence Index |
50 |
57 |
| Observed output speed |
Approximately 275 tokens/s |
Approximately 36 tokens/s |
| Input price |
$1.50/M |
$3/M |
| Output price |
$7.50/M |
$15/M |
| Context window |
Approximately 1M |
Approximately 1M |
| Model type |
Proprietary |
Open-weight release announced |
| Best fit |
Fast and cost-efficient agents |
Higher-intelligence, long-horizon work |
Kimi K3 has the stronger overall Intelligence Index score. It is a better candidate when maximum reasoning quality and sustained knowledge work matter more than response speed.
Gemini 3.6 Flash is approximately twice as cheap by listed input and output token prices, and its observed generation speed is several times higher. That makes it more attractive for interactive tools, coding loops, document analysis, and applications serving many simultaneous users.
There is also an openness difference, although it requires a date-sensitive qualification. Moonshot described Kimi K3 as an open-source model, but its official launch documentation said the complete weights would be released by July 27, 2026. They were not yet fully available at this article’s July 22 cutoff.
The practical verdict is:
- Choose Gemini 3.6 Flash for speed, lower API cost, multimodal input, and Google-native tools.
- Choose Kimi K3 for stronger general intelligence, long-running knowledge work, and future self-hosting possibilities.
- Test both on the actual workflow before assuming a higher benchmark score will produce a lower cost per successful task.
You can also try Kimi K3 on GPT Proto to compare its behavior with other current models.
Availability and API Access
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available through:
- Google AI Studio
- Gemini API
- Gemini Enterprise Agent Platform
- Gemini app
Gemini 3.6 Flash is also available in Google Antigravity and the Gemini Enterprise app. Gemini 3.5 Flash-Lite is being rolled out in Google Search.
The official API model IDs are:
gemini-3.6-flash
gemini-3.5-flash-lite
Developers evaluating the upgrade can currently try Gemini 3.5 Flash on GPT Proto or browse the GPT Proto AI model collection to compare more than 200 text, image, video, and audio models through one platform.
Migration Notes: Do Not Treat It as a Model-ID-Only Upgrade
Gemini 3.6 Flash may look like an obvious replacement for Gemini 3.5 Flash, but Google introduced API behavior changes that can affect existing applications.
Starting with Gemini 3.6 Flash and Gemini 3.5 Flash-Lite:
temperature is deprecated.
top_p is deprecated.
top_k is deprecated.
- Prefilled model turns are no longer supported.
- A request whose last non-empty turn is a model message can return an HTTP 400 error in future implementations.
Google currently says the deprecated sampling parameters are ignored, but future model generations may reject them. Developers should remove them rather than relying on the API to continue ignoring the values.
Before migrating production traffic, test:
- Structured JSON output
- Function and tool calls
- Multi-turn message history
- Prompts that previously depended on temperature
- Long-document accuracy
- Thinking-level behavior
- Token use and cost per completed task
- Retry and timeout handling
This is particularly important for agentic applications, where a small change in tool selection or reasoning length can significantly affect total cost.
Limitations to Know Before Switching
Both new models are capable, but they still have important limitations.
They only generate text
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite can analyze images, videos, audio, and PDFs, but they do not directly generate media. Image, video, speech, and music generation require separate models.
A 1M context window does not guarantee perfect recall
The context limit tells you how much content the API can accept, not whether every fact will receive equal attention. Important instructions and evidence should still be clearly structured, especially in very long prompts.
Thinking tokens count toward output billing
Increasing the thinking level can improve difficult tasks, but it also increases output consumption and latency. Flash-Lite at a high thinking level may behave very differently from its minimal default.
GA does not eliminate hallucinations
Both official model cards list hallucination, occasional slowness, and timeouts as known limitations. High-stakes outputs still require grounding, validation, or human review.
Computer Use documentation is currently inconsistent
Google’s launch guide states that the new models support its built-in tool suite, including Computer Use. However, the individual Gemini 3.5 Flash-Lite model page currently lists Computer Use as unsupported, while the launch materials describe it as available.
Developers planning browser or interface automation with Flash-Lite should verify the latest API capability documentation before production deployment.
Final Verdict
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are not redundant releases.
Gemini 3.6 Flash is the better default when a task requires coding, planning, complex reasoning, multimodal interpretation, or multiple tool calls. Its main advantage over Gemini 3.5 Flash is not a dramatic increase in general intelligence, but a better combination of speed, token efficiency, coding reliability, and output pricing.
Gemini 3.5 Flash-Lite is the better choice when the task is clearly defined and must be repeated thousands or millions of times. It costs more than the previous Flash-Lite generation in several pricing categories, but it delivers a significant improvement in reasoning, tool reliability, and throughput.
For many production systems, the best decision will be to use both: Gemini 3.6 Flash as the planner and Gemini 3.5 Flash-Lite as the scalable execution layer.
While GPT Proto prepares access to the two new endpoints, developers can explore GPT Proto and compare currently available Gemini, Kimi, GLM, DeepSeek, MiniMax, and other AI models through one API platform.