1M Token Massive Context
Ingest entire technical manuals or legal archives at once. The signature 1M context window allows for large-scale RAG and deep document analysis at a fraction of the price of larger models.

text
text
Explore the core technical strengths of the Gemini 3.5 Flash Lite API, from its massive memory to its industry-leading speed, designed specifically for high-efficiency enterprise deployments.
Ingest entire technical manuals or legal archives at once. The signature 1M context window allows for large-scale RAG and deep document analysis at a fraction of the price of larger models.

Process images, video up to one hour, audio, and PDFs natively. Gemini 3.5 Flash-Lite achieves a 68.2% score on MMMU benchmarks, outperforming many text-only lite competitors significantly.

Features an optimized internal head for schema adherence, resulting in a 98% success rate on complex nested extraction. Perfect for converting unstructured data into actionable insights.

Optimized for an instant-feel, the model delivers a Time To First Token approximately 30-40% faster than standard versions, ensuring smooth user experiences in chat and real-time tools.

Getting a gemini-3.5-flash-lite API key takes four steps and a few minutes. Create a free GPTProto account, add credits, generate your key, and make your first call — at $0.18 / $1.5 it's a cheaper gemini-3.5-flash-lite API key than going direct, and one key works across every model on the platform. Full gemini-3.5-flash-lite Documentation is in the docs.

Sign up

Top up

Generate your API key

Make your first API call
Get details on the Gemini 3.5 Flash Lite API. Learn about latency, the 1M context window, pricing, and how this model compares to others in the 3.5 family for building responsive AI applications.