GPT Proto

GPTProto

  • Dashboard
  • LLM

    • claude
      Claude Opus 5New
    • google
      Gemini 3.6 Flash
    • google
      Gemini 3.5 Flash Lite
    • moonshotai
      Kimi K3
    • openai
      GPT 5.6 Luna

    Image

    • bytedance
      Dola Seedream 5.0 Pro 260628New
    • google
      Gemini 3.1 Flash Lite Image
    • google
      Gemini 3.1 Flash Image
    • openai
      GPT Image 2
    • google
      Gemini 3.1 Flash Image Preview

    Video

    • kling
      Kling v3.0 4kNew
    • bytedance
      Dreamina Seedance 2.0 Mini 260615
    • kling
      Kling v3 Omni 4k
    • bytedance
      Dreamina Seedance 2.0 Fast 260128
    • bytedance
      Dreamina Seedance 2.0 260128
    Explore 214+ Models >
  • Generator

    • Create Image
    • Create Video
    • Edit in Canvas

    Features

    • Anime to Real Life AINew
    • Anime AI Art Generator
    • AI Object Remover
    • AI Image Editor
    • Unrestricted AI Image Generator
    • AI Motion Transfer
    • AI Clothes Remover
    • AI Watermark Remover
    • AI Image Enhancer Online
    • Online Background Remover Tool
    Explore All >

    Prompts

    • Seedance 2.0 PromptsNew
    • GPT Image 2 Prompts
    • Nano Banana Pro Prompts
    • Seedream 5.0 Pro Prompts
  • AI Blog

    • GLM 5.2 vs MiniMax M3: Which Is Better for Coding and Frontend Work?
    • How to Create Your Own AI Character With an API—No Coding Required
    • Kimi K3 vs Claude Opus 5: Which Is Better for Coding and AI Agents?
    • 20 Free Seedream 5.0 Pro Packaging Design Prompts for Products and E-commerce
    • GLM-5.2 vs Kimi K3 for Coding: Which Is Better for Developers in 2026?
    Explore All >

    AI Insight

    • What Is Emochi AI—and Why Is It Growing So Fast? (2026)
    • What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?
    • 12 Best AI Video Generation Tools in 2026 for YouTube, TikTok, Text and Images
    • What Is Qwen 3.8 Max? Release Date, 2.4T Preview, Pricing, and Early Benchmarks
    • Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?
    Explore All >

    AI Docs

    • gpt-image-2
    • gpt-5.4
    • kimi-k2.5
    • claude-opus-4-6
    • kling-v3.0-pro
    Explore All >

    AI Skills

    • browser-use
    • claude-to-im
    • competitive-ads-extractor
    • content-creator
    • data-storytelling
    Explore All >
Pricing
English繁體中文한국어日本語EspañolРусский
Get Started Now
  1. Home
  2. /Model
  3. /Google
  4. /gemini-3.6-flash
Google
gemini-3.6-flash
ChatDocumentation
Documentation
Access the Google Gemini 3.6 Flash API through GPTProto at $0.90 per 1M input tokens and $4.50 per 1M output tokens. The model accepts text, images, video, audio, and PDFs, with a 1,048,576-token input limit, 65,536-token output limit, configurable thinking, and agentic tools.

$ 0.9
$ 1.5

$ 4.5
$ 7.5

text

text

$ 0.9
$ 1.5

text

$ 4.5
$ 7.5

text

Related Models
All Models
Claude
Claude
claude-opus-5
$ 20
$ 25
Google
Google
gemini-3.5-flash-lite
$ 1.5
$ 2.5
MoonshotAI
MoonshotAI
kimi-k3
$ 13.5
$ 15
OpenAI
OpenAI
gpt-5.6-luna
$ 0.96
$ 1.2
Grok
Grok
grok-4.5
$ 3.6
$ 6
MiniMax
MiniMax
MiniMax-M3
$ 0.96
$ 1.2

Gemini 3.6 Flash API for Agentic Coding and Multimodal Workflows

Use Google’s generally available Flash model with 1M context, 64K output, four thinking levels, and native tools at 40% below Google’s standard token pricing on GPTProto.

1M Multimodal Context

Process text, images, video, audio, and PDF inputs within a 1,048,576-token window and return up to 65,536 text tokens. For video, 1M-context models support one hour at default media resolution or three hours at low resolution.

1M Context Diagram

More Efficient Agentic Coding

Gemini 3.6 Flash scored 49% on DeepSWE and 58.7% on SWE-Bench Pro in Google-reported evaluations. It makes fewer unwanted code edits and uses 17% fewer output tokens than Gemini 3.5 Flash on Artificial Analysis workloads.

gmini 3.6 Code View

Four Thinking Levels

Set thinking_level to minimal, low, medium, or high to balance latency, token use, and reasoning depth. Medium is the default; use lower levels for routine extraction and higher levels for multi-step coding or analysis.

Reasoning Toggle

Computer Use and Tool Support

Google lists function calling and computer use among the model’s tools. Gemini 3.6 Flash reached 83.0% on Google’s OSWorld-Verified evaluation, but computer use remains a preview capability.

gmini 3.6 RPA UI

What Is the Gemini 3.6 Flash API?

The Gemini 3.6 Flash API gives developers programmatic access to Google’s generally available Flash model for coding agents, multimodal analysis, long-context knowledge work, and tool-driven automation. Google released the stable gemini-3.6-flash model on July 21, 2026. Unlike the earlier gemini-3-flash-preview, this is a stable model ID intended for ongoing application use.

Gemini 3.6 Flash accepts text, images, video, audio, and PDF files, but it returns text only. Its 1,048,576-token input window can hold large repositories, long document sets, or extended media inputs, while the 65,536-token output limit supports detailed reports, code, and structured responses. Google lists caching, structured output, function calling, code execution, file search, URL context, Search grounding, Maps grounding, and configurable thinking among the underlying model’s capabilities.

The capability boundaries matter. Gemini 3.6 Flash does not generate images or audio, and it does not support the Live API. Computer use is available as a preview feature rather than a fully stable capability. When an integration depends on a Google-native tool, confirm that the GPTProto compatibility layer exposes it before migration. If your product requires speech-to-speech interaction, native image generation, or strict stability for desktop automation, route that part of the workload to another model.

Specification Gemini 3.6 Flash
Provider Google
Status Generally available
Stable model ID gemini-3.6-flash
Input types Text, image, video, audio, PDF
Output type Text
Input token limit 1,048,576
Output token limit 65,536
Thinking levels minimal, low, medium, high
Default thinking level medium
Google-listed tools Function calling, code execution, file search, Search, Maps, URL context
Computer use Supported in preview
Not supported Image generation, audio generation, Live API

Where Gemini 3.6 Flash Fits Best

Gemini 3.6 Flash is most useful when a task needs both a large working context and repeated tool calls. It is not automatically the cheapest choice for every short prompt, and its default medium thinking level can spend more tokens than a non-reasoning model on simple classification. Match the thinking level and model route to the actual job.

Workload Why it fits Implementation note
Repository coding agents Fewer unwanted edits and execution loops than Gemini 3.5 Flash in Google’s testing Use medium or high thinking and validate changes with tests
Multimodal document analysis Reads PDFs, images, charts, audio, and video in the same request Ask for structured output when downstream systems need stable fields
Long-context research 1M input tokens and supported context caching Cache repeated corpora instead of resending unchanged files
Browser or desktop automation Native computer-use tooling and 83.0% OSWorld-Verified in Google’s evaluation Treat computer use as preview and keep approval gates for consequential actions
Grounded assistants The underlying model supports Search, Maps, file search, and URL context Enable only the tools each request needs and track resulting token use

Gemini 3.6 Flash vs Gemini 3.5 Flash

Gemini 3.6 Flash is a direct efficiency upgrade over Gemini 3.5 Flash rather than a larger-context replacement. Both models keep the same 1M context class and default medium thinking level. The difference is task efficiency: Google reports 17% fewer output tokens on Artificial Analysis workloads, fewer reasoning turns and tool calls, and stronger results in coding, machine-learning engineering, computer use, and long-context retrieval.

The official input rate remains $1.50 per 1M tokens, while the official output rate falls from $9.00 to $7.50. On GPTProto, Gemini 3.6 Flash costs $0.90 per 1M input tokens and $4.50 per 1M output tokens; Gemini 3.5 Flash is currently $0.90 and $5.40. For new agentic or multimodal systems, 3.6 Flash is the better default. Keep 3.5 Flash only when you have already validated its behavior and need time to regression-test the migration.

Metric Gemini 3.6 Flash Gemini 3.5 Flash
Context window 1,048,576 1,048,576
Max output 65,536 65,536
Default thinking Medium Medium
Google standard input / output per 1M $1.50 / $7.50 $1.50 / $9.00
GPTProto input / output per 1M $0.90 / $4.50 $0.90 / $5.40
DeepSWE v1.1 49.0% 37.0%
MLE-Bench 63.9% 49.7%
OSWorld-Verified 83.0% 78.4%
GDM-MRCR v2 at 1M 54.0% 26.6%

Benchmark figures above are provider-reported. Use them to choose candidates, then run both models on your own prompts, tools, and success criteria before moving production traffic.

API Changes to Check Before You Migrate

Do not treat migration to 3.6 Flash as only a model-name replacement if your application uses Google’s native Interactions API. Google changed several generation and conversation fields for the latest models:

  • Change the model ID to gemini-3.6-flash.
  • Remove temperature, top_p, and top_k from generation configuration.
  • Replace thinking_budget with thinking_level; supported values are minimal, low, medium, and high.
  • Remove candidate_count, which Gemini 3.x does not support.
  • Do not send prefilled model turns. For multi-turn Interactions API sessions, use the server-side previous_interaction_id.
  • Regression-test tool schemas, structured outputs, and file handling before shifting all traffic.

If you use GPTProto’s OpenAI-compatible request format, follow the Quick Start and GPTProto documentation on this page rather than copying Google-native fields directly. The practical advantage is that one GPTProto API key and balance can also call Gemini 3.5 Flash, GPT-5.6 Luna, Claude Opus 4.8, Kimi K3, and 200+ other models, so you can run the same evaluation set without opening separate provider accounts.

How to Get a gemini-3.6-flash API Key

Getting a gemini-3.6-flash API key takes four steps and a few minutes. Create a free GPTProto account, add credits, generate your key, and make your first call — at $0.9 / $4.5 it's a cheaper gemini-3.6-flash API key than going direct, and one key works across every model on the platform. Full gemini-3.6-flash Documentation is in the docs.

Sign up

Sign up

Create your free GPT Proto account to begin. You can set up an organization for your team at any time.

Top up

Top up

Your balance can be used across all models on the platform, including gemini-3.6-flash, giving you the flexibility to experiment and scale as needed.

Generate your API key

Generate your API key

In your dashboard, create an API key — you'll need it to authenticate when making requests to gemini-3.6-flash.

Make your first API call

Make your first API call

Use your API key with our sample code to send a request to gemini-3.6-flash via GPT Proto and see instant AI-powered results.

Get API Key

Frequently Asked Questions

Is Gemini 3 Flash available via API?

Yes. Gemini 3.6 Flash is generally available through the Gemini API under the stable model ID gemini-3.6-flash, and GPTProto provides access through the model string shown on this page. Do not confuse it with gemini-3-flash-preview, which is an earlier preview model.

How much does the Gemini 3.6 Flash API cost on GPTProto?

GPTProto lists Gemini 3.6 Flash at $0.90 per 1M input tokens and $4.50 per 1M output tokens. Google’s standard rate is $1.50 input and $7.50 output, so GPTProto’s displayed token rates are 40% lower.

How do I get a Gemini 3.6 Flash API key?

Create a GPTProto account, add credit, generate a key in the dashboard, and use the gemini-3.6-flash model string from the Quick Start. The same key and balance work across more than 200 supported models, which makes fallback routing and A/B testing easier.

What inputs, context window, and output limit does Gemini 3.6 Flash support?

The model accepts text, images, video, audio, and PDFs. It supports up to 1,048,576 input tokens and 65,536 output tokens, and returns text. It supports structured output and tools, but it does not natively generate images or audio.

How does Gemini 3.6 Flash compare with Gemini 3.5 Flash?

The context and default thinking level are unchanged, but 3.6 Flash is more efficient. Google reports 17% fewer output tokens, DeepSWE improving from 37% to 49%, and OSWorld-Verified improving from 78.4% to 83.0%. Its official output price also falls from $9.00 to $7.50 per 1M tokens.

Gemini 3.6 Flash vs GPT-5.6, Claude Opus 4.8, or Kimi K3: which should I use?

Choose Gemini 3.6 Flash when you need Google’s multimodal inputs, Search or Maps grounding, computer-use tooling, and a 1M context window at a Flash-tier price. In Google’s published table, GPT-5.6 Luna leads 3.6 Flash on several coding benchmarks, while Gemini leads on OSWorld-Verified and chart reasoning. Google did not publish matched Opus 4.8 or Kimi K3 results in that table, so a universal winner claim would be misleading. Run the same task-level evaluation across the models with one GPTProto key.

Should I use Gemini 3.6 Flash instead of Gemini 3 Pro?

For a new API integration, Gemini 3.6 Flash has the clearer deployment path because it is generally available, lower cost, and optimized for repeated agentic and multimodal work. Gemini 3 Pro was a preview model with a different cost and latency profile. If an existing application depends on its output behavior, compare both on your own evaluation set before migrating.

GPT Proto

Empowering AI Innovation with Global Scale and Stability:

With our flagship product GPT Proto, we offer a unified interface to access and combine APIs from the world's leading AI providers—spanning text, vision, speech, and beyond. We empower developers and enterprises to simplify integration and accelerate innovation without limits.

Global Infrastructure, Local Compliance:

To ensure enterprise-grade reliability and compliance, Talent Tech Global Limited operates specifically as our global Billing and Contracting Entity. Meanwhile, our core technical infrastructure and R&D teams are strategically distributed across global innovation hubs, including Silicon Valley, Singapore, and Hong Kong.

Built to Scale:

We understand that stability is paramount. Our platform is built on a robust, decentralized architecture supporting dynamic Auto-scaling. Whether you are running a pilot or handling millions of concurrent requests, our system expands instantly to meet demand—guaranteeing that your business never outgrows our infrastructure.

Navigation

  • Dashboard
  • Models
  • Create Image
  • AI Image Upscale
  • AI Background Remover
  • Create Video
  • Edit in Canvas
  • Features
  • Pricing
  • AI Docs
  • AI Blog
  • AI Insight
  • AI Skills

Features

  • Anime to Real Life AI
  • Anime AI Art Generator
  • AI Object Remover
  • AI Image Editor
  • Unrestricted AI Image Generator
  • AI Motion Transfer
  • AI Clothes Remover
  • AI Watermark Remover
  • AI Image Enhancer Online
  • Online Background Remover Tool
  • AI Face Swap Image
  • AI Passport Photo Maker
  • MS Paint AI Generator
Explore all features >

LLM

  • Claude Opus 5
  • Gemini 3.6 Flash
  • Gemini 3.5 Flash Lite
  • Kimi K3
  • GPT 5.6 Luna
  • GPT 5.6 Terra
  • GPT 5.6 Sol
  • Grok 4.5
  • Claude Sonnet 5
  • Minimax M3
  • GLM 5.2
  • GPT 5.1 Chat Latest
  • Claude Fable 5
  • Qwen3.7 Max
  • Claude Opus 4.8 Thinking
  • Claude Opus 4.8
  • Gemini 3.5 Flash
  • DeepSeek v4 Flash
  • DeepSeek v4 Pro
  • Grok 4.3
Explore all models >

Image

  • Dola Seedream 5.0 Pro 260628
  • Gemini 3.1 Flash Lite Image
  • Gemini 3.1 Flash Image
  • GPT Image 2
  • Gemini 3.1 Flash Image Preview
  • Seedream 5.0 260128
  • Doubao Seedream 5.0 260128
  • Vidu Q2
  • Grok Imagine Image
  • Kling Image O1
  • GPT Image 1.5
  • Seedream 4.5 251128
  • Doubao Seedream 4.5 251128
  • Grok Imagine 0.9
  • Gemini 3 Pro Image Preview
  • Qwen Image Lora
  • Qwen Image Plus Lora
  • Qwen Image Plus
  • Grok 4 Image
  • GPT Image 1 Mini
Explore all models >

Video

  • Kling v3.0 4k
  • Dreamina Seedance 2.0 Mini 260615
  • Kling v3 Omni 4k
  • Dreamina Seedance 2.0 Fast 260128
  • Dreamina Seedance 2.0 260128
  • Vidu 2.0
  • Doubao Seedance 2.0 260128
  • Doubao Seedance 2.0 Fast 260128
  • Kling v3 Omni Pro
  • Kling v3 Omni Std
  • Vidu Q3 Turbo
  • Kling v3.0 Pro
  • Kling v3.0 Std
  • Vidu Q3 Pro
  • Kling v2.6 Std
  • Vidu Q2 Pro
  • Vidu Q2 Turbo
  • Vidu Q2 Pro Fast
  • Vidu Q2
  • Kling Video O1 Pro
Explore all models >

© 2026 Talent Tech Global Limited (Hong Kong) / Talent Tech Global LLC (US). All rights reserved.

  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap