Schuyler Stacy2026-07-23

Gemini 3.6 Flash y Gemini 3.5 Flash-Lite explicados: ¿cuál deberías usar?

Compara los precios, la velocidad, los benchmarks y los casos de uso de Gemini 3.6 Flash y Gemini 3.5 Flash-Lite. Descubre qué nuevo modelo de Google se adapta mejor a tu carga de trabajo de IA.

Gemini 3.6 Flash y Gemini 3.5 Flash-Lite explicados: ¿cuál deberías usar?

Resumen

  • Google lanzó Gemini 3.6 Flash y Gemini 3.5 Flash-Lite el 21 de julio de 2026. Ambos modelos están disponibles de forma general, no como versiones preliminares experimentales.
  • Gemini 3.6 Flash es la opción más potente para programación, análisis multimodal, trabajo de conocimiento y flujos de trabajo agénticos complejos.
  • Gemini 3.5 Flash-Lite es el modelo de ejecución más rápido y económico para extracción de documentos, traducción, clasificación, búsqueda y tareas de subagentes de gran volumen.
  • Ambos modelos admiten una ventana de contexto de 1 millón de tokens, hasta 64.000 tokens de salida, entrada multimodal, razonamiento, llamadas a funciones, salida estructurada y grounding mediante búsqueda.
  • Gemini 3.6 Flash no es considerablemente más inteligente que Gemini 3.5 Flash en todos los benchmarks. Sus principales ventajas son un menor número de tokens de salida, menos llamadas a herramientas, tiempos de tarea más cortos y un precio de salida inferior.
  • La forma más interesante de usarlos quizá no sea elegir un solo modelo. Puede consistir en utilizar Gemini 3.6 Flash como agente principal y Gemini 3.5 Flash-Lite como capa de ejecución.
  • Ambos modelos ya están disponibles mediante GPTProto con un 40 % de descuento sobre los precios de lista de Google. Puedes llamar a la API de Gemini 3.6 Flash o a la API de Gemini 3.5 Flash-Lite con una sola clave de API de GPTProto y un endpoint compatible con OpenAI.

Google ha lanzado dos nuevos modelos Flash listos para producción, pero sus nombres no explican de inmediato en qué se diferencian.

¿Gemini 3.5 Flash-Lite es simplemente un Gemini 3.6 Flash más pequeño? ¿“GA” indica un modelo distinto? ¿Gemini 3.6 Flash es realmente mejor que Gemini 3.5 Flash o solo es más barato?

La respuesta breve es que Google ha creado dos modelos para dos capas diferentes de un sistema de IA. Gemini 3.6 Flash está diseñado para tomar decisiones más difíciles y coordinar flujos de trabajo complejos. Gemini 3.5 Flash-Lite está diseñado para ejecutar grandes cantidades de tareas pequeñas de forma rápida y económica.

Esta guía cubre todo lo que necesitas saber sobre Gemini 3.6 Flash y Gemini 3.5 Flash-Lite, incluidas sus fechas de lanzamiento, precios, especificaciones, casos de uso, requisitos de migración y la comparación de Gemini 3.6 Flash con Gemini 3.5 Flash y Kimi K3.

Tabla de contenido

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite at a Glance

Google released both models on July 21, 2026. Unlike many recent Gemini launches, neither model entered the API as a temporary preview. Google lists both as stable and generally available for production use.

Specification Gemini 3.6 Flash Gemini 3.5 Flash-Lite
Release date July 21, 2026 July 21, 2026
Availability GA / Stable GA / Stable
Model ID gemini-3.6-flash gemini-3.5-flash-lite
Primary role General-purpose workhorse High-throughput execution model
Default thinking level Medium Minimal
Input context limit 1,048,576 tokens 1,048,576 tokens
Maximum output 65,536 tokens 65,536 tokens
Input types Text, image, video, audio, PDF Text, image, video, audio, PDF
Output type Text Text
Google standard input price $1.50 per 1M tokens $0.30 per 1M tokens
Google standard output price $7.50 per 1M tokens $2.50 per 1M tokens
GPT Proto input price $0.90 per 1M tokens $0.18 per 1M tokens
GPT Proto output price $4.50 per 1M tokens $1.50 per 1M tokens
GPT Proto discount 40% off 40% off
Best suited for Coding, complex agents, knowledge work, spatial and multimodal reasoning Extraction, classification, translation, search, data processing, subagents

Both models have a March 2026 knowledge cutoff, according to their official model cards. For more recent information, developers need to use search grounding or provide current data in the prompt.

What Is Gemini 3.6 Flash?

Gemini 3.6 Flash is Google’s latest general-purpose Flash model for coding, knowledge work, multimodal analysis, and multi-step agentic execution.

It is based on Gemini 3.5 Flash rather than being an entirely separate generation of the Gemini architecture. The main goal of the update is to make Flash more efficient during real work: fewer unnecessary output tokens, fewer reasoning steps, fewer tool calls, and fewer repeated execution loops.

That distinction matters. Gemini 3.6 Flash is not simply trying to produce a higher score on every intelligence benchmark. It is trying to complete the same or better work with less computation and less waiting.

According to Google’s launch announcement, Gemini 3.6 Flash uses approximately 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index. Google also reports that it takes fewer turns and tool calls to complete multi-step workflows.

Official evaluations show improvements in several practical areas:

  • DeepSWE increased from 37% to 49%, suggesting fewer incorrect code edits and execution loops.
  • MLE-Bench increased from 49.7% to 63.9% for machine learning research tasks.
  • OSWorld-Verified increased from 78.4% to 83% for computer-use tasks.
  • GDPval-AA v2 increased from 1349 to 1421 for knowledge-work performance.

These are vendor-reported results, so they should not be treated as universal guarantees. However, they support Google’s positioning of Gemini 3.6 Flash as a more reliable model for code migrations, technical diagnostics, document analysis, chart interpretation, and tool-using agents.

The model accepts text, images, video, audio, and PDF files, but it only generates text. It does not directly generate images, videos, or speech.

What Is Gemini 3.5 Flash-Lite—and What Does GA Mean?

Gemini 3.5 Flash-Lite is Google’s fastest and least expensive model in the Gemini 3.5 family. It is optimized for low-latency, high-volume tasks where throughput and API cost matter more than maximum reasoning quality.

The “GA” in Flash-Lite GA means Generally Available. It is an availability status, not part of the model’s name. A GA model is intended for stable production use rather than short-term preview testing.

Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite—not Gemini 3.6 Flash. Therefore, it should not be understood as a compressed version of the new 3.6 model. The two releases come from different upgrade paths:

  • Gemini 3.5 Flash → Gemini 3.6 Flash
  • Gemini 3.1 Flash-Lite → Gemini 3.5 Flash-Lite

Google positions Flash-Lite for use cases such as:

  • Document classification and structured extraction
  • Receipt and invoice processing
  • Product attribute extraction
  • Translation and localization
  • Search-result processing
  • High-volume customer-support routing
  • Tabular data processing
  • Repetitive tool calls
  • Parallel subagent execution

Independent testing cited in Google’s announcement measured Gemini 3.5 Flash-Lite at approximately 350 output tokens per second. Actual API performance will vary with prompt length, thinking level, server load, tools, and region, but the result illustrates the model’s throughput-focused design.

Flash-Lite is also considerably more capable than its predecessor. Google reports improvements from 31% to 54% on Terminal-Bench 2.1 and from 60.1% to 72.2% on its long-context evaluation.

It can use higher thinking levels for more complicated tasks, but doing so changes its cost and latency profile. If every request requires deep reasoning, Flash-Lite may lose some of the economic advantage that makes it attractive.

What Do “Flash” and “Flash-Lite” Mean in Gemini?

“Flash” is not an official acronym. It is Google’s product label for Gemini models that prioritize speed, efficiency, and scalable inference.

The practical meaning of Flash is:

  • Faster than heavier flagship models
  • Less expensive to run at scale
  • Suitable for interactive applications
  • Still capable of reasoning, coding, and multimodal understanding

Flash-Lite moves further toward the efficiency end of that spectrum. It is designed for workloads with a large number of relatively bounded tasks, such as extracting fields from thousands of documents or translating millions of short content items.

“Lite” does not mean that the model cannot reason or process multimodal inputs. Gemini 3.5 Flash-Lite still supports thinking, function calling, structured output, search grounding, and a 1M-token context window.

The difference is how Google expects each model to be deployed:

  • Flash: Choose it when the model must understand a complicated goal and decide what to do.
  • Flash-Lite: Choose it when the task is already defined and needs to be completed quickly at scale.

Gemini 3.6 Flash vs Gemini 3.5 Flash: What Actually Improved?

Gemini 3.6 Flash is the direct successor to Gemini 3.5 Flash, but “successor” does not mean that every intelligence score has increased.

The most meaningful improvements are efficiency, coding reliability, tool use, and total time per task.

Comparison Gemini 3.6 Flash Gemini 3.5 Flash
Standard input price $1.50/M $1.50/M
Standard output price $7.50/M $9/M
Default thinking level Medium Medium
Context window 1M 1M
Maximum output 64K 64K
Artificial Analysis Intelligence Index 50 50
Average time per task in launch testing 1.3 minutes 2.7 minutes
Primary advantage Lower token use and faster task completion Established production baseline

In Artificial Analysis testing, both models scored 50 on its Intelligence Index. That makes it difficult to argue that Gemini 3.6 Flash represents a large increase in general intelligence.

However, Gemini 3.6 Flash completed the tested tasks in less than half the average time. Its output price is also approximately 16.7% lower, while the model uses fewer output tokens in many agentic workloads.

For production users, this may be more valuable than a small benchmark increase. A model that reaches a similar answer with fewer tokens, fewer failed tool calls, and fewer correction loops can reduce both the infrastructure cost and the time users spend waiting.

There is little pricing incentive to begin a new deployment on Gemini 3.5 Flash when Gemini 3.6 Flash has the same standard input price and a lower output price. Existing applications, however, should still test migration compatibility instead of changing the model ID without review.

Developers who are not ready to migrate can continue to test Gemini 3.5 Flash on GPT Proto before comparing its real output quality, latency, and token consumption with Gemini 3.6 Flash.

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Pricing

The following prices are Google’s paid Gemini API rates per one million tokens at launch.

Model Standard Input Standard Output Batch Input Batch Output
Gemini 3.6 Flash $1.50 $7.50 $0.75 $3.75
Gemini 3.5 Flash $1.50 $9.00 $0.75 $4.50
Gemini 3.5 Flash-Lite $0.30 $2.50 $0.15 $1.25
Gemini 3.1 Flash-Lite* $0.25 $1.50 $0.125 $0.75

*Gemini 3.1 Flash-Lite charges a higher input rate for audio. Refer to the current Gemini API pricing page before deploying a production workload.

A Simple Cost Example

Suppose a workload consumes 100 million input tokens and generates 20 million output tokens.

Ignoring caching, grounding, and storage fees, the standard Google API cost would be:

  • Gemini 3.6 Flash: $150 input + $150 output = $300
  • Gemini 3.5 Flash: $150 input + $180 output = $330
  • Gemini 3.5 Flash-Lite: $30 input + $50 output = $80

This illustrates the large sticker-price difference between Flash and Flash-Lite. However, fixed token calculations do not tell the entire story.

A cheaper model can become expensive if it requires more retries, produces longer reasoning traces, or fails to complete a task. Similarly, a model with a higher token rate may have a lower cost per completed task if it uses fewer turns and tools.

That is why Gemini 3.6 Flash’s token efficiency matters. If it produces approximately 17% fewer output tokens while also charging 16.7% less for each output token, the savings on the output portion of some workloads can be substantially greater than the price-table difference alone suggests.

Gemini 3.5 Flash-Lite presents the opposite lesson. It is the least expensive model in the 3.5 family, but it is not cheaper than 3.1 Flash-Lite for every token type. Its text, image, and video input price increased from $0.25 to $0.30 per million tokens, while output increased from $1.50 to $2.50.

Developers are paying more than they did for 3.1 Flash-Lite, but they are receiving much better reasoning, agentic performance, and throughput.

GPT Proto Pricing: Access Both Models at 40% Off

Both models are also available through GPT Proto at 40% below Google’s listed standard token rates.

Model Google Input GPT Proto Input Google Output GPT Proto Output
Gemini 3.6 Flash $1.50/M $0.90/M $7.50/M $4.50/M
Gemini 3.5 Flash-Lite $0.30/M $0.18/M $2.50/M $1.50/M

At these rates, the same workload of 100 million input tokens and 20 million output tokens would cost approximately:

  • Gemini 3.6 Flash on GPT Proto: $90 input + $90 output = $180
  • Gemini 3.5 Flash-Lite on GPT Proto: $18 input + $30 output = $48

The same workload would cost approximately $300 and $80 respectively at Google’s standard listed rates.

GPT Proto uses pay-as-you-go billing, so developers do not need a separate subscription for each model. The same API key and account balance can also be used across GPT Proto’s collection of 200+ text, image, video, and audio models.

You can check the current rates on the Gemini 3.6 Flash API page and Gemini 3.5 Flash-Lite API page.

Which Gemini Flash Model Should You Choose?

The best model depends on whether your bottleneck is task complexity or execution volume.

Workload Recommended model Why
Large codebase changes Gemini 3.6 Flash Better planning and fewer unwanted edits
Multi-step coding agent Gemini 3.6 Flash Stronger tool use and execution loops
Chart and document analysis Gemini 3.6 Flash Better multimodal and spatial reasoning
Complex business research Gemini 3.6 Flash Stronger knowledge-work performance
Document classification Gemini 3.5 Flash-Lite Lower cost and higher throughput
Receipt or invoice extraction Gemini 3.5 Flash-Lite Designed for structured document processing
Large-scale translation Gemini 3.5 Flash-Lite Fast, multimodal, and inexpensive
Search-result processing Gemini 3.5 Flash-Lite Suitable for parallel high-volume execution
Simple subagent tasks Gemini 3.5 Flash-Lite Low-cost execution with adjustable thinking
Agent orchestration Both Flash plans; Flash-Lite executes

The More Interesting Answer: Use Both

Imagine an e-commerce platform that needs to extract product information from thousands of listings, PDFs, images, and supplier documents.

Gemini 3.6 Flash could:

  1. Inspect several representative documents.
  2. Design the extraction schema.
  3. Decide which sources are reliable.
  4. Identify ambiguous or conflicting product information.
  5. Review exceptions that require deeper reasoning.

Gemini 3.5 Flash-Lite could then run in parallel to:

  1. Read every product document.
  2. Extract brand, material, size, price, and availability.
  3. Translate product descriptions.
  4. Return structured JSON.
  5. Flag unusual records for the main agent.

This architecture avoids paying 3.6 Flash rates for every repetitive extraction while still using the stronger model where its reasoning matters.

Gemini 3.6 Flash vs Kimi K3

Gemini 3.6 Flash and Kimi K3 were released within days of each other, but they optimize for different things.

Kimi K3 is Moonshot AI’s 2.8-trillion-parameter flagship model for long-horizon coding, reasoning, and end-to-end knowledge work. Gemini 3.6 Flash emphasizes speed, token efficiency, multimodal workflows, and integration with Google’s tools.

The following figures come from the same Artificial Analysis comparison, making them more useful than comparing unrelated vendor benchmark tables.

Comparison Gemini 3.6 Flash Kimi K3
Intelligence Index 50 57
Observed output speed Approximately 275 tokens/s Approximately 36 tokens/s
Input price $1.50/M $3/M
Output price $7.50/M $15/M
Context window Approximately 1M Approximately 1M
Model type Proprietary Open-weight release announced
Best fit Fast and cost-efficient agents Higher-intelligence, long-horizon work

Kimi K3 has the stronger overall Intelligence Index score. It is a better candidate when maximum reasoning quality and sustained knowledge work matter more than response speed.

Gemini 3.6 Flash is approximately twice as cheap by listed input and output token prices, and its observed generation speed is several times higher. That makes it more attractive for interactive tools, coding loops, document analysis, and applications serving many simultaneous users.

There is also an openness difference, although it requires a date-sensitive qualification. Moonshot described Kimi K3 as an open-source model, but its official launch documentation said the complete weights would be released by July 27, 2026. They were not yet fully available at this article’s July 23 update.

The practical verdict is:

  • Choose Gemini 3.6 Flash for speed, lower API cost, multimodal input, and Google-native tools.
  • Choose Kimi K3 for stronger general intelligence, long-running knowledge work, and future self-hosting possibilities.
  • Test both on the actual workflow before assuming a higher benchmark score will produce a lower cost per successful task.

You can also try Kimi K3 on GPT Proto to compare its behavior with other current models.

Where Can You Access Gemini 3.6 Flash and Gemini 3.5 Flash-Lite?

Google provides direct access to the two models through Google AI Studio, the Gemini API, the Gemini app, and its enterprise products. Gemini 3.6 Flash is also available through Google Antigravity.

Developers who want one key for multiple AI providers can now access both models through GPT Proto:

GPT Proto provides an OpenAI-compatible API interface. Instead of maintaining separate Google, Kimi, OpenAI, DeepSeek, and other provider accounts, developers can use one API key, one balance, and the same base URL across 200+ models.

The GPT Proto model strings are:

gemini-3.6-flash
gemini-3.5-flash-lite

This also makes A/B testing easier. An application can route complex requests to Gemini 3.6 Flash and high-volume extraction or classification jobs to Gemini 3.5 Flash-Lite without rebuilding its API integration.

How to Use Gemini 3.6 Flash and Gemini 3.5 Flash-Lite API on GPT Proto

GPT Proto supports an OpenAI-compatible Chat Completions endpoint, so developers who already use the OpenAI SDK can access the two Gemini models by changing the API key, base URL, and model name.

Step 1: Create a GPT Proto API Key

Create or sign in to your GPT Proto account, add usage credits, and open the API Keys section in the dashboard.

Generate a new API key and store it securely. Do not paste the key directly into public repositories, frontend code, or shared documents.

For macOS or Linux, save it as an environment variable:

export GPTPROTO_API_KEY="sk-your-gptproto-api-key"

Step 2: Install the OpenAI Python SDK

Install or update the OpenAI SDK:

python -m pip install openai

GPT Proto uses the following OpenAI-compatible base URL:

https://gptproto.com/v1

Step 3: Call the Gemini 3.6 Flash API

Use Gemini 3.6 Flash when the request involves coding, planning, multimodal analysis, knowledge work, or a complex sequence of decisions.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["GPTPROTO_API_KEY"],
    base_url="https://gptproto.com/v1",
)

response = client.chat.completions.create(
    model="gemini-3.6-flash",
    messages=[
        {
            "role": "system",
            "content": (
                "You are a careful software engineering assistant. "
                "Inspect the problem before proposing changes and explain "
                "how each recommendation should be verified."
            ),
        },
        {
            "role": "user",
            "content": (
                "Review this API migration plan and identify compatibility, "
                "security, and performance risks."
            ),
        },
    ],
    max_completion_tokens=2048,
)

print(response.choices[0].message.content)

This request is sent to GPT Proto’s unified endpoint while using gemini-3.6-flash as the model ID.

Step 4: Switch to Gemini 3.5 Flash-Lite

To use Flash-Lite, keep the same client, API key, and base URL. Only change the model string:

response = client.chat.completions.create(
    model="gemini-3.5-flash-lite",
    messages=[
        {
            "role": "user",
            "content": (
                "Extract the merchant, invoice date, currency, subtotal, tax, "
                "and total from the following invoice text. Return valid JSON."
            ),
        }
    ],
    max_completion_tokens=1024,
)

print(response.choices[0].message.content)

Gemini 3.5 Flash-Lite is the better option for large batches of extraction, translation, classification, routing, and other clearly defined tasks.

Step 5: Route Tasks Between the Two Models

A production application does not have to send every request to the same model. A simple routing rule can use:

  • gemini-3.6-flash for planning, coding, exception handling, and difficult multimodal analysis.
  • gemini-3.5-flash-lite for repetitive execution, structured extraction, translation, and high-volume processing.

Because both models use the same GPT Proto endpoint and account balance, switching between them only requires changing the model ID.

Avoid adding deprecated Gemini sampling parameters such as temperature, top_p, and top_k to new integrations. Google has deprecated these parameters for Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.

Migration Notes: Do Not Treat It as a Model-ID-Only Upgrade

Gemini 3.6 Flash may look like an obvious replacement for Gemini 3.5 Flash, but Google introduced API behavior changes that can affect existing applications.

Starting with Gemini 3.6 Flash and Gemini 3.5 Flash-Lite:

  • temperature is deprecated.
  • top_p is deprecated.
  • top_k is deprecated.
  • Prefilled model turns are no longer supported.
  • A request whose last non-empty turn is a model message can return an HTTP 400 error in future implementations.

Google currently says the deprecated sampling parameters are ignored, but future model generations may reject them. Developers should remove them rather than relying on the API to continue ignoring the values.

Before migrating production traffic, test:

  • Structured JSON output
  • Function and tool calls
  • Multi-turn message history
  • Prompts that previously depended on temperature
  • Long-document accuracy
  • Thinking-level behavior
  • Token use and cost per completed task
  • Retry and timeout handling

This is particularly important for agentic applications, where a small change in tool selection or reasoning length can significantly affect total cost.

Limitations to Know Before Switching

Both new models are capable, but they still have important limitations.

They Only Generate Text

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite can analyze images, videos, audio, and PDFs, but they do not directly generate media. Image, video, speech, and music generation require separate models.

A 1M Context Window Does Not Guarantee Perfect Recall

The context limit tells you how much content the API can accept, not whether every fact will receive equal attention. Important instructions and evidence should still be clearly structured, especially in very long prompts.

Thinking Tokens Count Toward Output Billing

Increasing the thinking level can improve difficult tasks, but it also increases output consumption and latency. Flash-Lite at a high thinking level may behave very differently from its minimal default.

GA Does Not Eliminate Hallucinations

Both official model cards list hallucination, occasional slowness, and timeouts as known limitations. High-stakes outputs still require grounding, validation, or human review.

Computer Use Documentation Is Currently Inconsistent

Google’s launch guide states that the new models support its built-in tool suite, including Computer Use. However, the individual Gemini 3.5 Flash-Lite model page currently lists Computer Use as unsupported, while the launch materials describe it as available.

Developers planning browser or interface automation with Flash-Lite should verify the latest API capability documentation before production deployment.

Final Verdict

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are not redundant releases.

Gemini 3.6 Flash is the better default when a task requires coding, planning, complex reasoning, multimodal interpretation, or multiple tool calls. Its main advantage over Gemini 3.5 Flash is not a dramatic increase in general intelligence, but a better combination of speed, token efficiency, coding reliability, and output pricing.

Gemini 3.5 Flash-Lite is the better choice when the task is clearly defined and must be repeated thousands or millions of times. It costs more than the previous Flash-Lite generation in several pricing categories, but it delivers a significant improvement in reasoning, tool reliability, and throughput.

For many production systems, the best decision will be to use both: Gemini 3.6 Flash as the planner and Gemini 3.5 Flash-Lite as the scalable execution layer.

Both models are now available on GPT Proto at 40% off Google’s standard listed rates. Developers can use the Gemini 3.6 Flash API for complex coding and agentic workflows, switch to the Gemini 3.5 Flash-Lite API for high-volume execution, or browse the full GPT Proto AI model collection to compare more than 200 models through one API platform.

Creative Studio

Genera imágenes, videos y más con APIs de producción.

Comenzar a crear
Creative Studio
Modelos relacionados
Todos los modelos
Google
40% OFF
Google
40% OFF
Google
40% OFF
MoonshotAI
10% OFF

Preguntas frecuentes

¿Cuándo se lanzaron Gemini 3.6 Flash y Gemini 3.5 Flash-Lite?

Google lanzó ambos modelos el 21 de julio de 2026. Estuvieron disponibles a través de Gemini API, Google AI Studio, la aplicación Gemini y los productos empresariales desde el mismo día.

¿Gemini 3.6 Flash está disponible de forma general?

Sí. Gemini 3.6 Flash está disponible de forma general y figura como un modelo de producción estable. Su ID de modelo de API es gemini-3.6-flash.

¿Qué significa GA en Gemini 3.5 Flash-Lite GA?

GA significa “Generally Available” (disponible de forma general). Indica que Gemini 3.5 Flash-Lite está listo para uso en producción y no es una versión preliminar temporal. “GA” es un estado de disponibilidad, no forma parte del nombre del modelo.

¿Cuánto cuestan Gemini 3.6 Flash y Gemini 3.5 Flash-Lite?

Gemini 3.6 Flash cuesta 1,50 $ por millón de tokens de entrada y 7,50 $ por millón de tokens de salida con la tarifa estándar de la API de pago. Gemini 3.5 Flash-Lite cuesta 0,30 $ por millón de tokens de entrada y 2,50 $ por millón de tokens de salida. El procesamiento por lotes ofrece tarifas de tokens aproximadamente un 50 % más bajas.

¿Gemini 3.6 Flash es mejor que Gemini 3.5 Flash?

Por lo general, Gemini 3.6 Flash es la mejor opción para producción porque tiene un precio de salida inferior, utiliza menos tokens de salida, completa las tareas más rápido y mejora en varios benchmarks de programación y capacidades agénticas. Sin embargo, ambos modelos obtuvieron la misma puntuación de 50 en el Intelligence Index de Artificial Analysis, por lo que la versión 3.6 es más una mejora de eficiencia y ejecución que un salto universal de inteligencia.

¿Cuál es la diferencia entre Flash y Flash-Lite?

Flash está diseñado para equilibrar inteligencia, velocidad y coste en tareas complejas del mundo real. Flash-Lite prioriza una menor latencia, un menor coste de API y un alto rendimiento para tareas repetitivas o acotadas, como extracción, traducción, clasificación y ejecución de subagentes.

¿Ambos modelos admiten una ventana de contexto de 1 millón de tokens?

Sí. Gemini 3.6 Flash y Gemini 3.5 Flash-Lite admiten hasta 1.048.576 tokens de entrada y hasta 65.536 tokens de salida.

¿Gemini 3.5 Flash-Lite puede procesar imágenes y vídeo?

Sí. Acepta texto, imágenes, vídeo, audio y archivos PDF como entrada. Sin embargo, su salida está limitada a texto.

¿Gemini 3.6 Flash y Flash-Lite son gratuitos?

Google ofrece un nivel gratuito limitado de Gemini API para usos elegibles. Las aplicaciones de producción que necesitan límites superiores, almacenamiento en caché, acceso a Batch API o controles avanzados normalmente requieren el nivel de pago. Las cuotas gratuitas y la disponibilidad regional pueden cambiar.

¿Deberían los desarrolladores migrar inmediatamente desde Gemini 3.5 Flash?

Gemini 3.6 Flash ofrece buenas ventajas de precio y eficiencia, pero los desarrolladores deberían probar primero sus prompts y flujos de trabajo con herramientas. Los parámetros de muestreo como temperature, top_p y top_k están obsoletos, y ya no se admiten turnos de modelo rellenados previamente.

¿Puedo usar Gemini 3.6 Flash y Gemini 3.5 Flash-Lite en GPTProto?

Sí. Ambos modelos están disponibles mediante la API compatible con OpenAI de GPTProto. Usa `gemini-3.6-flash` o `gemini-3.5-flash-lite` como ID de modelo con la URL base `https://gptproto.com/v1`. La misma clave de API y el mismo saldo de cuenta funcionan con los demás modelos disponibles en GPTProto.

¿Cuánto cuestan los nuevos modelos Gemini en GPTProto?

GPTProto ofrece ambos modelos con un 40 % de descuento sobre las tarifas estándar publicadas por Google. Gemini 3.6 Flash cuesta 0,90 $ por millón de tokens de entrada y 4,50 $ por millón de tokens de salida. Gemini 3.5 Flash-Lite cuesta 0,18 $ por millón de tokens de entrada y 1,50 $ por millón de tokens de salida. Consulta la página de cada modelo para conocer la tarifa actual antes de calcular los costes de producción.

Artículos relacionados

Más blogs
GLM-5.2 vs Kimi K3 para programar: ¿cuál es mejor para los desarrolladores en 2026?

GLM-5.2 vs Kimi K3 para programar: ¿cuál es mejor para los desarrolladores en 2026?

En resumen: Kimi K3 es el modelo de programación más potente cuando la tarea es difícil, prolongada o visual. Supera a GLM-5.2 en la comparación de programación publicada por Moonshot y acepta imágenes y vídeo mediante su servicio alojado. GLM-5.2 sigue siendo la mejor opción predeterminada para el trabajo rutinario en repositorios: cuesta mucho menos, es más pequeño de operar y utiliza la permisiva licencia MIT. Kimi K3 también ha publicado sus pesos, pero su repositorio de 1,56 TB, la implementación recomendada con más de 64 aceleradores y su licencia personalizada hacen que el autoalojamiento suponga un compromiso considerablemente mayor. Elige Kimi cuando la capacidad sea el factor limitante; elige GLM cuando el coste y la sencillez operativa sean importantes cada día. La parte interesante de la comparación de código GLM-5.2 frente a Kimi K3 no es que ambos modelos puedan escribir un componente de React o resolver un algoritmo corto. Los modelos de este nivel ya superan ese umbral. La pregunta útil es qué ocurre cuando la tarea se complica: una auditoría de un repositorio, una migración de varios archivos, un error que solo aparece en una captura de pantalla o un prototipo jugable de Three.js que debe mantener la coherencia entre varios sistemas. Ahí es también donde la diferencia de precio empieza a importar. Kimi K3 ofrece mejores resultados en las pruebas públicas más difíciles, pero su precio oficial de salida es más de tres veces superior al de GLM-5.2. Un equipo que ejecute miles de revisiones ordinarias puede realizar más trabajo por dólar con GLM. Un desarrollador que intente rescatar un proyecto visual difícil probablemente pagará con gusto por K3.

Tiffany Layne | 2026-07-28

Kimi K3 vs GPT-5.6 Sol: ¿Tokens más baratos o tareas más baratas?

Kimi K3 vs GPT-5.6 Sol: ¿Tokens más baratos o tareas más baratas?

TL;DR Actualización — 28 de julio de 2026 : Los pesos completos de Kimi K3 ya son públicos. Moonshot AI publicó el checkpoint de 2,8 T, el informe técnico y la licencia de Kimi K3 en sus repositorios oficiales. El lanzamiento refuerza el argumento a favor del control y el despliegue de K3 frente a GPT-5.6 Sol, pero no cambia los resultados de las pruebas independientes ni hace que operar K3 por cuenta propia sea barato. Kimi K3 es más barato por token. GPT-5.6 Sol es la opción predeterminada más sólida para agentes de producción de alto riesgo. Ambas afirmaciones pueden ser ciertas. La diferencia es menor de lo que sugieren las tarjetas de precios. En las pruebas de Artificial Analysis, GPT-5.6 Sol max obtiene 59 en el Intelligence Index, frente a 57 de Kimi K3. Sin embargo, el coste medido por tarea es de aproximadamente 1,04 $ para Sol y 0,95 $ para K3—no la diferencia de dos a uno que implican sus precios oficiales de salida. Mi respuesta breve: elige GPT-5.6 Sol cuando la fiabilidad general, el rendimiento de los agentes de programación y el conjunto de herramientas alojadas de OpenAI sean lo más importante. Elige Kimi K3 cuando la entrada de vídeo, el trabajo con contexto extenso, un precio de lista más bajo o el acceso a pesos abiertos publicados cambien la decisión.

Schuyler Stacy | 2026-07-28

¿Qué es Kimi K3 y está realmente cerca de GPT-5.6 y Fable 5?

¿Qué es Kimi K3 y está realmente cerca de GPT-5.6 y Fable 5?

TL;DR Kimi K3 es el modelo multimodal de Moonshot AI, con 2,8 billones de parámetros, diseñado para programación de largo recorrido, trabajo del conocimiento, razonamiento y flujos de trabajo con agentes. Las pruebas independientes lo sitúan cerca de Claude Opus 4.8 y GPT-5.5 en general, mientras que GPT-5.6 Sol y Claude Fable 5 siguen por delante. K3 se acerca más en los benchmarks agénticos y lidera algunas pruebas de automatización, pero su tasa medida de alucinaciones aumentó frente a K2.6. Kimi K3 ahora está disponible con pesos abiertos. Moonshot AI ha publicado el checkpoint completo, la ficha del modelo, el informe técnico y la licencia personalizada Kimi K3. El repositorio oficial de Hugging Face ocupa aproximadamente 1,56 TB distribuidos en 96 fragmentos safetensors, y Moonshot recomienda implementaciones en supernodos con 64 aceleradores o más. Los pesos abiertos resuelven la cuestión de la propiedad. No convierten a K3 en un modelo local convencional. Para la mayoría de los desarrolladores, la API alojada sigue siendo el punto de partida práctico. La API de Kimi K3 en GPTProto muestra actualmente un precio de 2,70 $ por millón de tokens de entrada y 13,50 $ por millón de tokens de salida. Elige los pesos cuando el control de los datos, la inferencia personalizada o la modificación del modelo justifiquen la infraestructura y la revisión de la licencia. En resumen, Kimi K3 está lo bastante cerca de GPT-5.6 y Fable 5 como para formar parte de la misma conversación—y su lanzamiento con pesos abiertos ofrece ahora a los desarrolladores una opción de implementación que ninguno de los dos modelos cerrados ofrece.

Michael Johnson | 2026-07-28

La mejor API de IA para desarrolladores en 2026: comparación de 10 plataformas

La mejor API de IA para desarrolladores en 2026: comparación de 10 plataformas

TL;DR Mejores API directas: OpenAI es la opción predeterminada más segura para uso general; Anthropic Claude destaca en programación y agentes de larga duración; Gemini es ideal para prototipos multimodales de bajo coste; y DeepSeek lidera en precio por token de texto. Mejores opciones multimodelo: OpenRouter es la opción más clara para probar muchos LLM. GPTProto es más adecuado cuando un producto necesita modelos de texto, imagen y vídeo bajo una sola clave de API y un saldo compartido. Mejores opciones de infraestructura: Amazon Bedrock encaja en implementaciones empresariales gobernadas por AWS, mientras que Replicate, fal.ai y Together AI son más adecuados para inferencia de modelos abiertos o medios generativos. No existe un ganador universal. Compara la adecuación a la carga de trabajo, la cobertura de modelos, las unidades reales de facturación, los controles de producción y el coste de migración. Los precios y la disponibilidad se comprobaron el 14 de julio de 2026; verifica las páginas actuales de los proveedores antes de implementar.

Tiffany Layne | 2026-07-15