Multimodaler Kontext mit 1 Mio. Tokens
Analysieren Sie Text, Bilder, Videos, Audio und PDFs innerhalb eines Eingabefensters mit 1.048.576 Tokens. Die Antworten bestehen ausschließlich aus Text und umfassen maximal 65.536 Tokens.
curl --request POST "https://gptproto.com/v1/chat/completions" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "gemini-3.8-flash",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}'Chat, Coding-Agenten und Dokumentenarbeit. Preise pro 1 Mio. Tokens – Eingabe, gecachte Eingabe und Ausgabe werden separat abgerechnet. GPTProto liegt 40% unter den offiziellen Preisen.
| Szenario | Google Liste | OpenRouter | GPTProto | Du sparst / Monat |
|---|---|---|---|---|
| Persönlich10 Mio. Tokens / Monat (4.8 Mio. im Cache) | $20.52 | $21.65 | $12.31 | −$8.21≈ $98.50 / Jahr |
| Team100 Mio. Tokens / Monat (48 Mio. im Cache) | $205.20 | $216.49 | $123.12 | −$82.08≈ $984.96 / Jahr |
| Business500 Mio. Tokens / Monat (240 Mio. im Cache) | $1026.00 | $1082.43 | $615.60 | −$410.40≈ $4924.80 / Jahr |
Führen Sie Gemini 3.8 Flash über GPTProto für Softwareaufgaben mit mehreren Dateien, iterative Tool-Nutzung und Multimedia-Analysen aus. Das Modell verarbeitet Text, Bilder, Videos, Audio und PDFs und gibt bis zu 65.536 Text-Tokens aus.
Multimodaler Kontext mit 1 Mio. Tokens
Analysieren Sie Text, Bilder, Videos, Audio und PDFs innerhalb eines Eingabefensters mit 1.048.576 Tokens. Die Antworten bestehen ausschließlich aus Text und umfassen maximal 65.536 Tokens.
Fortschritte bei Coding-Aufgaben mit langem Planungshorizont
Nutzen Sie Gemini 3.8 Flash für Refactorings über mehrere Dateien und terminalgesteuerte Aufgaben. Google gibt für Terminal-Bench 2.1 einen Wert von 90,8 % an, verglichen mit 81,6 % für Gemini 3.7 Flash.
Anpassbare Denkstufen
Legen Sie einen niedrigen, mittleren oder hohen Denkaufwand fest, um Latenz, Tokenverbrauch und Tiefe der Schlussfolgerungen auszubalancieren. Mittel ist die Standardeinstellung; die minimale Stufe wird nicht unterstützt und führt zu einem Fehler.
Tools und strukturierte Ausgabe
Erstellen Sie Agenten mit Funktionsaufrufen, Codeausführung, Dateisuche, Google-Suche mit Grounding, URL-Kontext, Caching und strukturierter Ausgabe. Die Computernutzung wird ebenfalls als Vorschau unterstützt.
The Google Gemini 3.8 API provides programmatic access to Gemini 3.8 Flash, a generally available reasoning model released on September 2, 2026. Its official model ID is gemini-3.8-flash. Google positions it for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Gemini 3.8 is a multimodal-input, text-output model. A request can combine text, images, video, audio, or PDFs inside a 1,048,576-token input window; responses can contain up to 65,536 text tokens. Supported tools include function calling, code execution, file search, Google grounding, URL context, caching, structured output, and preview computer use. It does not generate images or audio and does not support the Live API.
On GPTProto, one balance and API key work across Gemini 3.8 Flash and 200+ other models, simplifying routing, fallback, and controlled A/B tests.
| Specification | Gemini 3.8 Flash |
|---|---|
| Provider | |
| Release status | Generally available (GA) |
| Release date | September 2, 2026 |
| Official model ID | gemini-3.8-flash |
| Input types | Text, image, video, audio, PDF |
| Output type | Text |
| Input context limit | 1,048,576 tokens |
| Maximum output | 65,536 tokens |
| Thinking levels | Low, medium, high; medium by default |
| Core tools | Function calling, code execution, file search, Search and Maps grounding, URL context |
| Other capabilities | Caching, structured output, Batch API, Flex inference, Priority inference |
Repository-scale coding: Provide source files, tests, issue context, and tools for multi-file debugging, refactoring, migration planning, and terminal-based validation.
Long-horizon agents: Use function calls for search, code execution, retrieval, and validation loops. Set explicit stopping conditions, tool permissions, and acceptance checks around the model.
Multimodal analysis: Combine PDFs, screenshots, charts, audio, and video with instructions for document review, meeting analysis, visual QA, or structured extraction.
Grounded workflows: Pair Google Search, Maps, file search, or URL context with structured output for research, internal knowledge tools, or location-aware analysis.
Gemini 3.8 keeps the same 1M-token context, 64K output limit, input types, and introductory Google token rates as Gemini 3.7 Flash. Google reports stronger coding, multimodal, and agent results, but 3.8 can use more tokens as it reasons, calls tools, and verifies work.
| Decision factor | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Terminal-Bench 2.1 | 90.8% | 81.6% |
| SWE-Bench Pro | 61.6% | 60.4% |
| SWE-Atlas | 51.9% | 48.0% |
| tau3-bench Banking | 38.1% | 30.9% |
| CharXiv multimodal | 86.2% | 84.5% |
| Humanity's Last Exam | 45.4% | 45.7% |
| Official introductory input/output rate | $0.75 / $3.75 per 1M | $0.75 / $3.75 per 1M |
| Best fit | Difficult coding, agents, mixed-media reasoning | Efficiency-first workflows with lower token use |
These are Google-reported evaluations, not independent GPTProto tests. Gemini 3.8 improves most listed results but not every row, and equal token rates do not guarantee equal cost per completed task.
Changing the model name is only the first migration step. Check these documented request rules:
Replace integer thinking_budget settings with thinking_level: low, medium, or high.
Do not send minimal; Gemini 3.8 Flash does not support it.
Remove frequency_penalty, presence_penalty, and candidate_count, which can return validation errors.
Do not rely on temperature, top_k, or top_p; Google documents these sampling fields as ignored for this model.
Match each function response to the preceding function call's ID, name, and execution count.
Remove prefilled model turns; do not end history with a model-role message.
For an OpenAI-compatible integration, keep the payload minimal and verify mapped fields in staging. Run a canary with the same prompts, tools, timeouts, and success criteria as the current model.
Cross-vendor benchmarks rarely use identical tools or budgets. Test each candidate on the same task and measure completion rate, invalid tool calls, latency, output tokens, and cost per successful run.
| Comparison query | Most useful test |
|---|---|
| Claude Fable 5.1 vs Gemini 3.8 | Multi-file implementation with tools, tests, and a fixed output budget |
| Gemini 3.8 vs GPT-5.6 | Code generation, debugging, structured output, and recovery from a failed tool call |
| Gemini 3.8 vs Grok 4.6 | Search-grounded research with source quality, latency, and cost measured separately |
| Qwen3.8-Max-0902vs Gemini 3.8 | Long-context coding and enterprise document analysis on the same input set |
| Gemini 3.8 vs Kimi K3 | Long agent loops with identical stopping rules and tool schemas |
| GLM-5.3 Flash vs Gemini 3.8 | High-volume coding tasks, total tokens used, and cost per accepted result |
Gemini 3.8 is a strong candidate when one request must combine mixed-media input with Google grounding or code tools. There is no universal winner without workload-level testing.
Choose Gemini 3.8 Flash for mixed-media input, a 1M-token context, long tool loops, or stronger terminal behavior than Gemini 3.7. It fits coding agents, multimodal research, document analysis, and structured enterprise workflows.
Keep Gemini 3.7 or a lower-cost alternative when token efficiency matters more than extra verification. For simple extraction or chat, test low thinking effort and compare cost per successful task.
Anleitungen, Vergleiche und Updates zu diesem Modell.
Alle Artikel
Claude Fable 5.1 und Mythos 5.1 basieren auf demselben Modell, unterscheiden sich aber beim Zugriff. Erfahre mehr über Preise, Funktionen, Benchmarks, Schutzmaßnahmen und Änderungen bei der API-Migration.

Qwen3.8-Flash-Next und GLM-5.3 Flash im Vergleich: Coding, Agenten, Frontend-Arbeit, Geschwindigkeit, Preise, Kontext, Lizenzen und produktiver Einsatz.

Was ist die Tencent Hy4 Preview? Erfahre mehr über den Veröffentlichungsstatus, das 770B-MoE-Design, den 1M-Kontext, API-Preise, Benchmarks, Limits und Modellvergleiche.

Verwalte deinen Tech-Stack effizienter mit einem API-Schlüssel für mehrere KI-Modelle. Greife über einen einzigen Endpunkt auf GPT, Claude und Gemini zu und spare Entwicklungszeit. Probiere es jetzt aus.