1M-Token Multimodal Context
Analyze text, images, video, audio, and PDFs within a 1,048,576-token input window. Responses are text-only, with a maximum output of 65,536 tokens.
curl --request POST "https://gptproto.com/v1/chat/completions" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "gemini-3.8-flash",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}'Estimate a request with real work scenarios. GPTProto token pricing is 40% below official rates.
Пополнение $100 даёт вам:
Кредиты за пополнение действуют бессрочно. Вы получите всего $100.00.
Дополнительная скидка 40% на модель, экономия $66.6649 по сравнению с официальным API Google.
Run Gemini 3.8 Flash through GPTProto for multi-file software work, iterative tool use, and mixed-media analysis. The model accepts text, images, video, audio, and PDFs while returning up to 65,536 text tokens.
Analyze text, images, video, audio, and PDFs within a 1,048,576-token input window. Responses are text-only, with a maximum output of 65,536 tokens.
Use Gemini 3.8 Flash for multi-file refactoring and terminal-driven work. Google reports 90.8% on Terminal-Bench 2.1, compared with 81.6% for Gemini 3.7 Flash.
Set low, medium, or high thinking effort to balance latency, token use, and reasoning depth. Medium is the default; the minimal level is unsupported and returns an error.
Build agents with function calling, code execution, file search, Google Search grounding, URL context, caching, and structured output. Computer use is also supported in preview.
The Google Gemini 3.8 API provides programmatic access to Gemini 3.8 Flash, a generally available reasoning model released on September 2, 2026. Its official model ID is gemini-3.8-flash. Google positions it for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Gemini 3.8 is a multimodal-input, text-output model. A request can combine text, images, video, audio, or PDFs inside a 1,048,576-token input window; responses can contain up to 65,536 text tokens. Supported tools include function calling, code execution, file search, Google grounding, URL context, caching, structured output, and preview computer use. It does not generate images or audio and does not support the Live API.
On GPTProto, one balance and API key work across Gemini 3.8 Flash and 200+ other models, simplifying routing, fallback, and controlled A/B tests.
| Specification | Gemini 3.8 Flash |
|---|---|
| Provider | |
| Release status | Generally available (GA) |
| Release date | September 2, 2026 |
| Official model ID | gemini-3.8-flash |
| Input types | Text, image, video, audio, PDF |
| Output type | Text |
| Input context limit | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Thinking levels | Low, medium, high; medium by default |
| Core tools | Function calling, code execution, file search, Search and Maps grounding, URL context |
| Other capabilities | Caching, structured output, Batch API, Flex inference, Priority inference |
Repository-scale coding: Provide source files, tests, issue context, and tools for multi-file debugging, refactoring, migration planning, and terminal-based validation.
Long-horizon agents: Use function calls for search, code execution, retrieval, and validation loops. Set explicit stopping conditions, tool permissions, and acceptance checks around the model.
Multimodal analysis: Combine PDFs, screenshots, charts, audio, and video with instructions for document review, meeting analysis, visual QA, or structured extraction.
Grounded workflows: Pair Google Search, Maps, file search, or URL context with structured output for research, internal knowledge tools, or location-aware analysis.
Gemini 3.8 keeps the same 1M-token context, 64K output limit, input types, and introductory Google token rates as Gemini 3.7 Flash. Google reports stronger coding, multimodal, and agent results, but 3.8 can use more tokens as it reasons, calls tools, and verifies work.
| Decision factor | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Terminal-Bench 2.1 | 90.8% | 81.6% |
| SWE-Bench Pro | 61.6% | 60.4% |
| SWE-Atlas | 51.9% | 48.0% |
| tau3-bench Banking | 38.1% | 30.9% |
| CharXiv multimodal | 86.2% | 84.5% |
| Humanity's Last Exam | 45.4% | 45.7% |
| Official introductory input/output rate | $0.75 / $3.75 per 1M | $0.75 / $3.75 per 1M |
| Best fit | Difficult coding, agents, mixed-media reasoning | Efficiency-first workflows with lower token use |
These are Google-reported evaluations, not independent GPTProto tests. Gemini 3.8 improves most listed results but not every row, and equal token rates do not guarantee equal cost per completed task.
Changing the model name is only the first migration step. Check these documented request rules:
Replace integer thinking_budget settings with thinking_level: low, medium, or high.
Do not send minimal; Gemini 3.8 Flash does not support it.
Remove frequency_penalty, presence_penalty, and candidate_count, which can return validation errors.
Do not rely on temperature, top_k, or top_p; Google documents these sampling fields as ignored for this model.
Match each function response to the preceding function call's ID, name, and execution count.
Remove prefilled model turns; do not end history with a model-role message.
For an OpenAI-compatible integration, keep the payload minimal and verify mapped fields in staging. Run a canary with the same prompts, tools, timeouts, and success criteria as the current model.
Cross-vendor benchmarks rarely use identical tools or budgets. Test each candidate on the same task and measure completion rate, invalid tool calls, latency, output tokens, and cost per successful run.
| Comparison query | Most useful test |
|---|---|
| Claude Fable 5.1 vs Gemini 3.8 | Multi-file implementation with tools, tests, and a fixed output budget |
| Gemini 3.8 vs GPT-5.6 | Code generation, debugging, structured output, and recovery from a failed tool call |
| Gemini 3.8 vs Grok 4.6 | Search-grounded research with source quality, latency, and cost measured separately |
| Qwen3.8-Max-0902vs Gemini 3.8 | Long-context coding and enterprise document analysis on the same input set |
| Gemini 3.8 vs Kimi K3 | Long agent loops with identical stopping rules and tool schemas |
| GLM-5.3 Flash vs Gemini 3.8 | High-volume coding tasks, total tokens used, and cost per accepted result |
Gemini 3.8 is a strong candidate when one request must combine mixed-media input with Google grounding or code tools. There is no universal winner without workload-level testing.
Choose Gemini 3.8 Flash for mixed-media input, a 1M-token context, long tool loops, or stronger terminal behavior than Gemini 3.7. It fits coding agents, multimodal research, document analysis, and structured enterprise workflows.
Keep Gemini 3.7 or a lower-cost alternative when token efficiency matters more than extra verification. For simple extraction or chat, test low thinking effort and compare cost per successful task.
Руководства, сравнения и обновления по этой модели.
Все статьи
Claude Fable 5.1 and Mythos 5.1 share one model but differ in access. See pricing, features, benchmarks, safeguards, and API migration changes.

Qwen3.8-Flash-Next vs GLM-5.3 Flash compared for coding, agents, frontend work, speed, pricing, context, licenses, and production use.

What is Tencent Hy4 Preview? See its release status, 770B MoE design, 1M context, API pricing, benchmarks, limits, and model comparisons.

Manage your stack better with one API key for multiple AI models. Access GPT, Claude, and Gemini from one endpoint and cut your dev time. Try it now.