Mis à jour le 7 août 2026 : Alibaba a officiellement lancé la version de production de Qwen3.8-Max le 3 août, remplaçant l’ancien modèle qwen3.8-max-preview comme modèle phare actuel de son API. Cet article a été mis à jour avec l’architecture confirmée, la fenêtre de contexte, les tarifs, la disponibilité de l’API et le calendrier des poids ouverts.
Qwen3.8-Max est le modèle Qwen le plus performant d’Alibaba à ce jour : un modèle multimodal Sparse Mixture-of-Experts doté de 2,4 billions de paramètres au total et de 95 milliards de paramètres actifs par requête.
Le modèle stable prend en charge les entrées texte, image et vidéo, produit des sorties textuelles et offre une fenêtre de contexte allant jusqu’à 1 million de tokens, avec jusqu’à 128 k tokens en sortie. Il est conçu pour le codage complexe, l’analyse visuelle, la recherche, les travaux professionnels et les tâches d’agents à long horizon.
Ce lancement change également la réponse pratique pour les développeurs. Qwen3.8-Max n’est plus limité à un aperçu susceptible d’évoluer ni à un forfait personnel exclusivement basé sur les crédits. Il dispose désormais d’une API normale à la consommation, Alibaba l’affichant à 2 $ par million de tokens d’entrée et 6 $ par million de tokens de sortie. Il est également disponible via l’API Qwen 3.8 Max sur GPTProto.
En bref : Qwen 3.8 Max est un aperçu réel et particulièrement intéressant, et non un lancement de produit finalisé. Les développeurs devraient le tester, noter la date de chaque résultat et éviter de planifier des migrations en production autour de spécifications qu’Alibaba n’a pas encore publiées.
Qwen 3.8 Max at a Glance
| Specification |
Confirmed Qwen3.8-Max Details |
| Developer |
Alibaba Qwen |
| Stable model name |
qwen3.8-max |
| Preview release |
July 19, 2026 |
| Production release |
August 3, 2026 |
| Architecture |
Sparse Mixture-of-Experts with hybrid attention |
| Total parameters |
2.4 trillion |
| Active parameters |
95 billion |
| Context window |
Up to 1 million tokens |
| Maximum output |
Up to 128K tokens |
| Inputs |
Text, images, and video |
| Output |
Text |
| Thinking mode |
Hybrid thinking with adjustable reasoning effort |
| Function calling |
Supported |
| Structured output |
Supported |
| Web search |
Supported |
| Context caching |
Supported |
| Batch inference |
Not currently supported |
| Fine-tuning |
Not currently supported |
| Official API price |
$2/M input and $6/M output tokens |
| GPT Proto availability |
Available now |
| Open weights |
Announced but not downloadable as of August 7, 2026 |
That list is shorter than the specification tables appearing in some early Qwen 3.8 Max coverage. It is shorter for a reason. A preview model does not gain a 1M context window, a particular Mixture-of-Experts layout, or a release license just because those details would look plausible beside other 2026 flagships.
Is Qwen 3.8 Max Released?
Yes. Qwen3.8-Max received its production release on August 3, 2026.
Alibaba first made qwen3.8-max-preview available on July 19. That preview could change between calls and did not have ordinary pay-as-you-go pricing. The stable qwen3.8-max release replaces that launch-week situation with a documented API model, published specifications, and standard per-token billing.
Developers can now call the model through Alibaba Cloud Model Studio or use the Qwen3.8-Max API on GPT Proto alongside other text, image, and video models under the same account.
What Does the 2.4T Parameter Count Mean?
Qwen3.8-Max contains 2.4 trillion total parameters, but its Sparse Mixture-of-Experts architecture activates approximately 95 billion parameters for each request.
That distinction matters. The full parameter count represents the model’s total capacity, while the active count more directly affects inference computation. About 4% of the total parameters are active at a time, allowing Alibaba to scale model capacity without paying the computational cost of activating all 2.4 trillion parameters for every token.
Parameter count alone still does not prove that a model is faster, cheaper, or more accurate. Serving infrastructure, token efficiency, reasoning length, caching, and the number of retries all affect the real cost of a completed task.
Is Qwen 3.8 Max Open Source?
Not yet as of August 7, 2026.
Alibaba has confirmed that the Qwen3.8-2.4T-A95B weights will be released, making this the first Qwen Max-class model scheduled for an open-weight release. The ModelScope release page currently points to August 12, 2026.
Until the checkpoint and license are actually published, Qwen3.8-Max should be described as an API model with an announced open-weight release—not as an already downloadable open-source model.
Even after the weights arrive, a 2.4-trillion-parameter checkpoint will require serious storage, memory, networking, and inference infrastructure. For most developers, the hosted API will remain considerably easier than self-hosting.
Qwen 3.8 Max API Pricing
Alibaba’s published pay-as-you-go rate is:
| Billing Item |
Official List Price |
| Input tokens |
$2 per 1M tokens |
| Output tokens |
$6 per 1M tokens |
| Context window |
Up to 1M tokens |
| Maximum output |
Up to 128K tokens |
This makes Qwen3.8-Max cheaper at official list price than Kimi K3’s $3/$15 input-output rate, while remaining more expensive than some smaller text-only models.
GPT Proto pricing may differ from Alibaba’s direct list price. Check the live Qwen 3.8 Max API page before estimating a large production workload.
What do the early Qwen 3.8 Max benchmarks show?
The evidence is stronger than it was during the July preview, but it still needs context.
At launch, Alibaba reported that Qwen3.8-Max ranked fifth in Text Arena, second in Vision Arena, and fourth in Frontend Code Arena. Alibaba also published examples of long-horizon coding, visual application reconstruction, professional document analysis, and autonomous tool use.
These results establish Qwen3.8-Max as a serious frontier model. They do not guarantee that it will beat every alternative on a specific repository or agent workflow.
The earlier 269-file StackPerf comparison can remain useful, but it tested the pre-release Qwen3.8-Max Preview. Kimi K3 scored 83/100 and the Qwen preview scored 80/100, while Qwen completed all 44 tool calls successfully and received the stronger tool-use score. Treat that result as a dated behavioral signal—not as a final ranking of the stable August release.
One useful early test comes from a matched StackPerf comparison against Kimi K3. Both models inspected frozen copies of 269 files and had 60 minutes to produce a repository-level integration design with citations, migration steps, tests, and an evidence ledger. The reports were anonymized before review.
Kimi K3 scored 83/100 after factual penalties; Qwen 3.8 Max Preview scored 80/100. Qwen earned 9/10 for tool use versus Kimi's 8/10, completed all 44 tool calls successfully, and produced the longer report with fewer requests. Kimi finished sooner, used fewer tokens on its tested route, and handled revisions and regeneration more completely.
That is credible evidence for one kind of long-context software architecture work. It is not proof that Kimi K3 is generally better, or that Qwen is generally better. The models ran through different providers, and route latency, caching, launch traffic, and reasoning settings all affected the result.
The early community picture is similarly mixed. In the main Hacker News launch discussion, one tester reported that an animated SVG task took more than 10 minutes of reasoning, while another developer said the preview was working well for coding but was still too new to judge. These are anecdotes, not benchmarks. They do flag two things worth measuring in your own tests: time to a usable answer and task completion per token.
Qwen 3.8 Max vs Qwen 3.7 Max
Qwen 3.8 Max looks like an upgrade in capability scope, but Qwen 3.7 Max remains the safer API choice today.
| Dimension |
Qwen 3.8 Max Preview |
Qwen 3.7 Max |
| Status |
Continuously updated preview |
Existing API model |
| Published parameters |
2.4T total |
Not the main public selling point |
| Vision |
Alibaba says yes |
Text-only |
| Context |
Not published |
1M tokens |
| Max output |
Not published |
65,536 tokens |
| Weights |
Promised, not released |
Closed-weight |
| Pricing |
Token Plan Credits; no normal token rate |
GPT Proto price: $0.36 input / $1.44 output per 1M |
| Production backend use |
Individual plan forbids automated backend use |
Conventional metered API access |
I would not perform a blind model-string upgrade from Qwen 3.7 Max to the 3.8 preview. First check output stability, tool schemas, latency, reasoning-token use, and whether the final production model preserves the preview's behavior.
If you need a Qwen API for an application now, the existing Qwen 3.7 Max endpoint is the practical bridge:
curl "https://gptproto.com/v1/chat/completions" \
-H "Authorization: Bearer $GPTPROTO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.7-max",
"messages": [
{
"role": "user",
"content": "Review this architecture and list the three highest-risk assumptions."
}
]
}'
This calls Qwen 3.7 Max, not Qwen 3.8 Max. That distinction prevents a common launch-week failure: publishing code for an endpoint the platform does not carry.
Qwen 3.8 Max vs Kimi K3, GLM-5.2, and Fable 5
The honest comparison is asymmetric because Qwen 3.8 Max is missing several public specifications.
| Model |
Best current reason to test it |
Main trade-off |
| Qwen 3.8 Max Preview |
New 2.4T Qwen flagship with vision and strong early agent results |
Moving preview, unknown context and token price, no weights yet |
| Kimi K3 |
Published 2.8T MoE architecture, 1M multimodal context, conventional API access |
Always-on reasoning can increase latency and output cost |
| GLM-5.2 |
MIT open weights, known active parameters, 1M context, lower-cost agentic coding |
Text-only; total capability may trail the newest closed previews on some tasks |
| Fable 5 |
The reference point named in Qwen's own launch claim for long-horizon work |
Higher price, and Qwen has not published evidence for its comparison claim |
For Qwen 3.8 Max vs Kimi K3, the best public matched test currently gives Kimi a three-point lead while giving Qwen the better tool-use subscore. For Qwen 3.8 Max vs GLM-5.2, the decision is less about an unverified quality ranking and more about deployment: GLM has downloadable MIT-licensed weights today; Qwen does not. For Qwen 3.8 Max vs Fable 5, there is not enough public evidence to validate Alibaba's ranking.
See the full Qwen 3.8 Max vs Kimi K3 comparison
Should Developers Use Qwen 3.8 Max Now?
Yes, if the workload benefits from complex coding, visual understanding, long context, or multi-stage agent execution.
The production API, published token price, 1M-token context window, and stable model name remove most of the deployment objections that applied to the July preview. Developers can now evaluate it as a normal hosted model instead of an experimental subscription-only endpoint.
There are still boundaries:
Batch inference and fine-tuning are not currently supported.
Its 128K output limit can produce expensive responses when high reasoning effort is used indiscriminately.
The downloadable weights and final license are not yet available as of August 7.
Alibaba’s longest autonomous-work examples are vendor-run demonstrations, not guarantees for every application.
For hosted production use, start with the Qwen3.8-Max API on GPT Proto. For self-hosting, wait until the announced checkpoint and license are actually published.