GPT Proto

GPTProto

  • Dashboard
  • LLM

    • claude
      Claude Opus 5New
    • google
      Gemini 3.6 Flash
    • google
      Gemini 3.5 Flash Lite
    • moonshotai
      Kimi K3
    • openai
      GPT 5.6 Luna

    Image

    • bytedance
      Dola Seedream 5.0 Pro 260628New
    • google
      Gemini 3.1 Flash Lite Image
    • google
      Gemini 3.1 Flash Image
    • openai
      GPT Image 2
    • google
      Gemini 3.1 Flash Image Preview

    Video

    • kling
      Kling v3.0 4kNew
    • bytedance
      Dreamina Seedance 2.0 Mini 260615
    • kling
      Kling v3 Omni 4k
    • bytedance
      Dreamina Seedance 2.0 Fast 260128
    • bytedance
      Dreamina Seedance 2.0 260128
    Explore 214+ Models >
  • Generator

    • Create Image
    • Create Video
    • Edit in Canvas

    Features

    • Anime to Real Life AINew
    • Anime AI Art Generator
    • AI Object Remover
    • AI Image Editor
    • Unrestricted AI Image Generator
    • AI Motion Transfer
    • AI Clothes Remover
    • AI Watermark Remover
    • AI Image Enhancer Online
    • Online Background Remover Tool
    Explore All >

    Prompts

    • Seedance 2.0 PromptsNew
    • GPT Image 2 Prompts
    • Nano Banana Pro Prompts
    • Seedream 5.0 Pro Prompts
  • AI Blog

    • GLM 5.2 vs MiniMax M3: Which Is Better for Coding and Frontend Work?
    • How to Create Your Own AI Character With an API—No Coding Required
    • Kimi K3 vs Claude Opus 5: Which Is Better for Coding and AI Agents?
    • 20 Free Seedream 5.0 Pro Packaging Design Prompts for Products and E-commerce
    • GLM-5.2 vs Kimi K3 for Coding: Which Is Better for Developers in 2026?
    Explore All >

    AI Insight

    • What Is Emochi AI—and Why Is It Growing So Fast? (2026)
    • What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?
    • 12 Best AI Video Generation Tools in 2026 for YouTube, TikTok, Text and Images
    • What Is Qwen 3.8 Max? Release Date, 2.4T Preview, Pricing, and Early Benchmarks
    • Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?
    Explore All >

    AI Docs

    • gpt-image-2
    • gpt-5.4
    • kimi-k2.5
    • claude-opus-4-6
    • kling-v3.0-pro
    Explore All >

    AI Skills

    • browser-use
    • claude-to-im
    • competitive-ads-extractor
    • content-creator
    • data-storytelling
    Explore All >
Pricing
English繁體中文한국어日本語EspañolРусский
Get Started Now
  1. Home
  2. /Model
  3. /Google
  4. /gemini-3.5-flash-lite
Google
gemini-3.5-flash-lite
ChatDocumentation
Documentation
Gemini 3.5 Flash is a cost-optimized, ultra-low latency model by Google. It features a 1M token context window and native multimodal support for text, video, and audio tasks at scale.

$ 0.18
$ 0.3

$ 1.5
$ 2.5

text

text

$ 0.18
$ 0.3

text

$ 1.5
$ 2.5

text

Related Models
All Models
Claude
Claude
claude-opus-5
$ 20
$ 25
Google
Google
gemini-3.6-flash
$ 4.5
$ 7.5
MoonshotAI
MoonshotAI
kimi-k3
$ 13.5
$ 15
OpenAI
OpenAI
gpt-5.6-luna
$ 0.96
$ 1.2
Grok
Grok
grok-4.5
$ 3.6
$ 6
MiniMax
MiniMax
MiniMax-M3
$ 0.96
$ 1.2

Gemini 3.5 Flash-Lite API

Access Google’s fastest and most cost-efficient Gemini 3.5 model on GPTProto for $0.18 per 1M input tokens and $1.50 per 1M output tokens, 40% below Google’s standard API rates. Use one GPTProto key and balance across 200+ text, image, video, and audio models.

350 Output Tokens per Second

Google reports roughly 350 output tokens per second based on Artificial Analysis testing. Flash-Lite is designed for workloads where throughput and response time matter more than maximum reasoning depth.

Fast Gemini API

Multimodal Input, Text Output

Send text, images, video, audio, and PDF files. The model returns text, making it suitable for document parsing, media understanding, transcription analysis, summarization, and multimodal data extraction.

Multimodal Gemini

Structured Outputs and Tool Use

Use schema-constrained outputs and function calling for extraction, routing, and tool-based agents. Confirm route-specific availability for code execution, file search, URL context, and search grounding before deployment.

Gemini JSON Mode

1M Token Context Window

Process up to 1,048,576 input tokens and generate up to 65,536 output tokens. This capacity supports large document sets, long transcripts, reports, repositories, and detailed subagent handoffs.

Gemini Context

What Is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google’s general-availability efficiency model for high-volume agentic tasks, translation, document processing, classification, and structured extraction. Released on July 21, 2026, it is the fastest model in the Gemini 3.5 series and is positioned for applications where latency, throughput, and API cost are stricter constraints than maximum reasoning depth.

The Google Gemini 3.5 Flash-Lite API accepts text, images, video, audio, and PDF files within a 1,048,576-token input window. It produces text responses of up to 65,536 tokens and supports thinking, structured outputs, function calling, caching, code execution, file search, URL context, and search grounding in Google’s native API. It does not generate images or audio, and it does not support the Live API. Availability of provider-specific tools can differ through an OpenAI-compatible gateway, so check the GPTProto documentation before relying on a Google-native extension.

Specification Gemini 3.5 Flash-Lite
Provider Google
Stable model ID gemini-3.5-flash-lite
Launch stage General availability
Release date July 21, 2026
Input limit 1,048,576 tokens
Maximum output 65,536 tokens
Accepted inputs Text, image, video, audio, PDF
Output Text
Core capabilities Thinking, structured outputs, function calling, caching
Not supported Image generation, audio generation, Live API, tuning

Best Uses for Gemini 3.5 Flash-Lite

Flash-Lite is most useful when an application repeats focused tasks at high volume. It can serve as a lower-cost worker model under a more capable orchestrator, or handle complete workflows whose decisions are narrow and easy to validate.

  • Classification and routing: Label support tickets, detect intent, assign documents, or choose the next tool without paying for a frontier model on every request.

  • Document parsing and JSON extraction: Read PDFs, invoices, receipts, reports, and product records, then return fields that match a defined schema.

  • Translation and summarization: Process large queues of multilingual text, meeting transcripts, reviews, or catalog content where unit cost affects total operating spend.

  • Multimodal review: Extract text and facts from images, recorded audio, video, and PDFs while keeping the output in a text-based downstream workflow.

  • Focused subagents: Delegate search, retrieval, code edits, verification, and tool calls to specialized workers while reserving a stronger model for planning or final review.

The model is less suitable when one difficult request requires the highest possible coding accuracy, long-horizon planning, or repeated recovery from tool errors. For those workloads, test Gemini 3.5 Flash or Gemini 3.6 Flash against the same evaluation set before choosing by token price alone.

Choose a Thinking Level by Task Complexity

Gemini 3.5 Flash-Lite supports configurable thinking in Google’s native API. Google recommends MINIMAL for latency-sensitive classification, routing, and JSON extraction. This setting is the default and limits unnecessary reasoning tokens on predictable tasks.

Use MEDIUM or HIGH for subagents that write code, run terminal commands, call several APIs, or must recover from intermediate errors. More thinking can improve multi-step execution, but it also increases output-token usage and response time. If you call the model through GPTProto’s OpenAI-compatible endpoint, check the current documentation for the mapped thinking parameter before sending provider-specific fields.

Gemini 3.5 Flash-Lite vs Gemini 3.5 Flash

Gemini 3.5 Flash-Lite and Gemini 3.5 Flash share a 1,048,576-token context window and 65,536-token maximum output, but they are designed for different workload economics. Flash-Lite prioritizes throughput and low unit cost; Flash targets more complex coding, agentic planning, and knowledge work.

Comparison Gemini 3.5 Flash-Lite Gemini 3.5 Flash
Best fit High-volume subagents, extraction, translation, document parsing Complex coding, agent planning, deeper knowledge work
Model positioning Fastest, lowest-cost model in the 3.5 series More capable Flash-tier model
Input context 1,048,576 tokens 1,048,576 tokens
Maximum output 65,536 tokens 65,536 tokens
Google standard input price $0.30 per 1M tokens $1.50 per 1M tokens
Google standard output price $2.50 per 1M tokens $9.00 per 1M tokens
GPTProto Flash-Lite price $0.18 input / $1.50 output per 1M tokens See the Gemini 3.5 Flash model page

At Google’s standard rates, Flash-Lite costs 80% less for input and about 72% less for output than Gemini 3.5 Flash. Choose it when request volume and response time dominate the cost of occasional retries. Choose Flash when a smaller number of harder tasks makes first-pass quality more important than the lowest token rate.

Using Gemini 3.5 Flash-Lite Through GPTProto

GPTProto provides Gemini 3.5 Flash-Lite API access through the same account, API key, and balance used for more than 200 models. This reduces account and billing fragmentation when an application routes simple extraction to Flash-Lite, complex reasoning to Gemini 3.5 Flash or Gemini 3.6 Flash, and media work to separate image, video, or audio models.

For an existing OpenAI-compatible application, keep the surrounding client structure and use the endpoint, authentication method, and model string shown in the GPTProto documentation. Do not assume that every Google-native field maps one-to-one. Test thinking controls, caching, grounding, file handling, tool schemas, and streaming behavior before moving production traffic.

Migration Notes for Existing Gemini Workloads

Google documents several compatibility differences that can affect migrations from older Gemini models. Custom temperature, top-K, and top-P values may be ignored. Custom frequency and presence penalties can return an error. Requests whose final conversation turn has the model role are also unsupported.

Before switching a high-volume workflow, replay a representative test set and check JSON-schema compliance, tool-call completion, token usage, latency, and retry frequency. A cost-effective Gemini 3.5 Flash-Lite API deployment should be evaluated by cost per successful task, not only cost per token. This is especially important for multi-step agents, where an overly low thinking setting can stop a tool sequence too early.

How to Get a gemini-3.5-flash-lite API Key

Getting a gemini-3.5-flash-lite API key takes four steps and a few minutes. Create a free GPTProto account, add credits, generate your key, and make your first call — at $0.18 / $1.5 it's a cheaper gemini-3.5-flash-lite API key than going direct, and one key works across every model on the platform. Full gemini-3.5-flash-lite Documentation is in the docs.

Sign up

Sign up

Create your free GPT Proto account to begin. You can set up an organization for your team at any time.

Top up

Top up

Your balance can be used across all models on the platform, including gemini-3.5-flash-lite, giving you the flexibility to experiment and scale as needed.

Generate your API key

Generate your API key

In your dashboard, create an API key — you'll need it to authenticate when making requests to gemini-3.5-flash-lite.

Make your first API call

Make your first API call

Use your API key with our sample code to send a request to gemini-3.5-flash-lite via GPT Proto and see instant AI-powered results.

Get API Key

Frequently Asked Questions

How much does the Gemini 3.5 Flash-Lite API cost on GPTProto?

GPTProto lists Gemini 3.5 Flash-Lite at $0.18 per 1M input tokens and $1.50 per 1M output tokens. Google’s standard API rates are $0.30 and $2.50 respectively, so the GPTProto rates are 40% lower. A request using 10,000 input tokens and 2,000 output tokens would cost approximately $0.0048.

Do I need a separate Google Gemini 3.5 Flash-Lite API key?

No. Gemini 3.5 Flash-Lite API access on GPTProto uses your GPTProto API key and shared account balance. The same key can call 200+ supported models, so you do not need a separate provider account or credit balance for each model.

What is the Gemini 3.5 Flash-Lite context limit?

The model accepts up to 1,048,576 input tokens and can return up to 65,536 output tokens. The limit applies across the combined prompt and multimodal inputs after tokenization, so large PDFs, audio, and video can consume the context differently from plain text.

Can Gemini 3.5 Flash-Lite process images, video, audio, and PDFs?

Yes. It accepts text, image, video, audio, and PDF inputs, then returns text. It can analyze or extract information from those media types, but it is not an image-generation, video-generation, text-to-speech, or real-time voice model.

Does Gemini 3.5 Flash-Lite support JSON mode and function calling?

Google lists structured outputs and function calling as supported capabilities. They are useful for schema-based extraction and tool-driven agents. Because gateway parameter mappings can differ, validate the exact request schema and tool behavior through GPTProto before deploying a critical workflow.

Should I use Gemini 3.5 Flash-Lite or Gemini 3.5 Flash?

Use Flash-Lite for high-volume classification, extraction, translation, document processing, and focused subagents. Use Gemini 3.5 Flash when complex coding, deeper planning, or stronger first-pass reasoning justifies a higher token cost. Test both on representative tasks when failure or retry cost is significant.

Is Gemini 3.5 Flash-Lite a stable model?

Yes. Google lists gemini-3.5-flash-lite as a general-availability model released on July 21, 2026. This distinguishes it from preview model IDs that may change more frequently or have tighter usage limits.

Related Articles

More Blogs
7 Best Unrestricted AI Image Generators in 2026 (Compared)

7 Best Unrestricted AI Image Generators in 2026 (Compared)

Stop hitting safety walls. Find the best unrestricted ai image generator to reclaim your creative freedom. Compare top tools and start creating today.

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?

Compare Gemini 3.6 Flash and Gemini 3.5 Flash-Lite pricing, speed, benchmarks and use cases. See which new Google model fits your AI workload.

Qwen 3.8 Max vs GLM 5.2: Which Is Better for Coding in 2026?

Qwen 3.8 Max vs GLM 5.2: Which Is Better for Coding in 2026?

Compare Qwen 3.8 Max vs GLM 5.2 on coding, API access, context, pricing, and open weights. See which model is safer for production in 2026.

What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?

What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?

Is Kimi K3 open source—and truly close to GPT-5.6 and Fable 5? Explore its 1M context, API pricing, independent benchmarks, and Reddit reaction.

GPT Proto

Empowering AI Innovation with Global Scale and Stability:

With our flagship product GPT Proto, we offer a unified interface to access and combine APIs from the world's leading AI providers—spanning text, vision, speech, and beyond. We empower developers and enterprises to simplify integration and accelerate innovation without limits.

Global Infrastructure, Local Compliance:

To ensure enterprise-grade reliability and compliance, Talent Tech Global Limited operates specifically as our global Billing and Contracting Entity. Meanwhile, our core technical infrastructure and R&D teams are strategically distributed across global innovation hubs, including Silicon Valley, Singapore, and Hong Kong.

Built to Scale:

We understand that stability is paramount. Our platform is built on a robust, decentralized architecture supporting dynamic Auto-scaling. Whether you are running a pilot or handling millions of concurrent requests, our system expands instantly to meet demand—guaranteeing that your business never outgrows our infrastructure.

Navigation

  • Dashboard
  • Models
  • Create Image
  • AI Image Upscale
  • AI Background Remover
  • Create Video
  • Edit in Canvas
  • Features
  • Pricing
  • AI Docs
  • AI Blog
  • AI Insight
  • AI Skills

Features

  • Anime to Real Life AI
  • Anime AI Art Generator
  • AI Object Remover
  • AI Image Editor
  • Unrestricted AI Image Generator
  • AI Motion Transfer
  • AI Clothes Remover
  • AI Watermark Remover
  • AI Image Enhancer Online
  • Online Background Remover Tool
  • AI Face Swap Image
  • AI Passport Photo Maker
  • MS Paint AI Generator
Explore all features >

LLM

  • Claude Opus 5
  • Gemini 3.6 Flash
  • Gemini 3.5 Flash Lite
  • Kimi K3
  • GPT 5.6 Luna
  • GPT 5.6 Terra
  • GPT 5.6 Sol
  • Grok 4.5
  • Claude Sonnet 5
  • Minimax M3
  • GLM 5.2
  • GPT 5.1 Chat Latest
  • Claude Fable 5
  • Qwen3.7 Max
  • Claude Opus 4.8 Thinking
  • Claude Opus 4.8
  • Gemini 3.5 Flash
  • DeepSeek v4 Flash
  • DeepSeek v4 Pro
  • Grok 4.3
Explore all models >

Image

  • Dola Seedream 5.0 Pro 260628
  • Gemini 3.1 Flash Lite Image
  • Gemini 3.1 Flash Image
  • GPT Image 2
  • Gemini 3.1 Flash Image Preview
  • Seedream 5.0 260128
  • Doubao Seedream 5.0 260128
  • Vidu Q2
  • Grok Imagine Image
  • Kling Image O1
  • GPT Image 1.5
  • Seedream 4.5 251128
  • Doubao Seedream 4.5 251128
  • Grok Imagine 0.9
  • Gemini 3 Pro Image Preview
  • Qwen Image Lora
  • Qwen Image Plus Lora
  • Qwen Image Plus
  • Grok 4 Image
  • GPT Image 1 Mini
Explore all models >

Video

  • Kling v3.0 4k
  • Dreamina Seedance 2.0 Mini 260615
  • Kling v3 Omni 4k
  • Dreamina Seedance 2.0 Fast 260128
  • Dreamina Seedance 2.0 260128
  • Vidu 2.0
  • Doubao Seedance 2.0 260128
  • Doubao Seedance 2.0 Fast 260128
  • Kling v3 Omni Pro
  • Kling v3 Omni Std
  • Vidu Q3 Turbo
  • Kling v3.0 Pro
  • Kling v3.0 Std
  • Vidu Q3 Pro
  • Kling v2.6 Std
  • Vidu Q2 Pro
  • Vidu Q2 Turbo
  • Vidu Q2 Pro Fast
  • Vidu Q2
  • Kling Video O1 Pro
Explore all models >

© 2026 Talent Tech Global Limited (Hong Kong) / Talent Tech Global LLC (US). All rights reserved.

  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap