Coding and Cowork Upgrade
Further post-training targets complex software engineering, scientific research, professional work, and multi-stage tasks that must continue from planning through delivery.
Estimate a request with real work scenarios. GPTProto token pricing is 10% below official rates.
Top-up $100 and you get:
Top-up credits with permanent validity. You will receive a total of $100.00.
Additional 10% model discount, saving $11.0902 versus direct official Qwen API calls.
Use Qwen3.8-Max-0902 for repository-scale coding, scientific research, professional knowledge work, and long-horizon agents. The revision retains the Qwen3.8-Max family's 1M-token context and multimodal input while adding further post-training for coding and cowork tasks.
Further post-training targets complex software engineering, scientific research, professional work, and multi-stage tasks that must continue from planning through delivery.
Process large repositories, long documents, tool history, and mixed source material within a 1,000,000-token context window, with up to 131,072 output tokens.
Combine text, images, and video for interface review, document analysis, chart reasoning, and recorded workflow inspection. The model returns text rather than generated media.
Connect external functions, request schema-constrained output, stream responses, and reuse cached context in agent workflows. Provider-managed tool availability can vary by endpoint.
Qwen3.8-Max-0902 is Alibaba Qwen's September 2, 2026 revision of Qwen3.8-Max. It is not a newly announced architecture or a larger successor. Qwen describes it as the existing 2.4-trillion-parameter mixture-of-experts model after additional post-training for coding and cowork, with the goal of improving complex enterprise tasks, scientific research, and long-horizon workflows.
The hosted model retains the Qwen3.8-Max family's documented operating envelope: text, image, and video input; text output; function calling; structured outputs; context caching; and thinking or non-thinking operation. Alibaba lists a nominal 1,000,000-token context window, with up to 991,808 input tokens in standard mode, 983,616 input tokens in thinking mode, and 131,072 output tokens.
The page name and request identifier require a careful distinction. Qwen announced the revision as Qwen3.8-Max-0902, while Alibaba's public Model Studio documentation still lists the general API model ID as qwen3.8-max. An API provider may expose the update through the standard alias or a dated snapshot. For GPTProto calls, use the exact model string displayed in Quick Start instead of constructing one from the page title.
| Specification | Confirmed Detail |
|---|---|
| Provider | Alibaba Qwen |
| Revision date | September 2, 2026 |
| Update type | Further post-training for coding and cowork |
| Published family size | 2.4T total parameters, mixture-of-experts architecture |
| Input modalities | Text, image, and video |
| Output modality | Text |
| Context window | 1,000,000 tokens |
| Maximum standard input | 991,808 tokens |
| Maximum thinking-mode input | 983,616 tokens |
| Maximum output | 131,072 tokens |
| Function calling | Supported |
| Structured outputs | Supported |
| Context caching | Supported |
| Fine-tuning | Not currently supported |
| Open-weight status | No separately labeled 0902 checkpoint has been identified |
The 0902 release is a behavioral upgrade rather than a specification reset. The parameter count, 1M context window, output ceiling, and supported input modalities remain within the published Qwen3.8-Max family envelope. The main change is further post-training aimed at better coding, cowork, tool use, and sustained task completion.
| Decision Factor | Original Qwen3.8-Max | Qwen3.8-Max-0902 |
|---|---|---|
| Model foundation | 2.4T MoE flagship | Same published family foundation |
| Context and output | 1M context; up to 131,072 output tokens | Unchanged published limits |
| Training focus | Coding, professional work, research, and agents | Additional post-training for coding and cowork |
| Code Arena: WebDev | 1,669 in Qwen's reported comparison | 1,691 at launch, a 22-point increase |
| Public API naming | Alibaba documents qwen3.8-max |
Use the provider's displayed alias or snapshot ID |
| Open weights | Separate Qwen3.8 family checkpoints exist | No separate 0902-labeled checkpoint confirmed |
The Code Arena result is a useful signal for frontend and agentic coding, but it does not prove that every repository or tool workflow will improve. Test the old and new routes with the same prompts, repository state, tool permissions, reasoning settings, and acceptance criteria before switching production traffic.
Inspect module boundaries, connect requirements to existing code, plan multi-file changes, call development tools, and review test output. The 1M context can hold source, configuration, documentation, and agent history, but executable tests and checkpoints are still required.
Use Qwen3.8-Max-0902 for migration agents, technical investigations, data-analysis pipelines, and internal assistants that query several tools before returning structured results. Validate tool arguments and restrict permissions at the application layer.
Combine papers, reports, tables, diagrams, and earlier findings in one task. Require source identifiers and a clear separation between extracted facts and inference. Use retrieval when the source set exceeds the context limit or changes frequently.
Analyze screenshots, charts, visual documents, and recorded workflows through text, image, and video input. The endpoint returns text, so media creation still requires a dedicated image or video model.
Compare cost per accepted result, including reasoning tokens, retries, failed tool calls, and human correction—not token price alone.
| Model | Start Here When | Validate Before Routing More Traffic |
|---|---|---|
| Qwen3.8-Max-0902 | The task combines difficult coding, long context, multimodal evidence, and many tool steps | Regression behavior, latency, and total tokens on your own workflow |
| GPT-5.6 | Your application already depends on GPT-oriented coding and agent behavior | The exact model tier, tool compatibility, latency, and cost |
| Gemini 3.7 Flash | High-volume multimodal work makes speed and token cost the first filter | Complex repository accuracy and required output length |
| Kimi K3 | Long coding or research tasks need another frontier route for A/B testing | First-pass completion, tool-call accuracy, and token use |
| GLM-5.3 Flash | Budget-sensitive coding and agent traffic needs a lower-cost first route | The hardest multi-step and repository-scale tasks |
For cost-sensitive routing, begin with a Flash-class model such as GLM-5.3 Flash, then send difficult or repeatedly failing tasks to Qwen3.8-Max-0902.
Before moving traffic from the original Qwen3.8 Max, record the model string, prompt, reasoning settings, tool schemas, latency, token use, and result. Run the same evaluation against 0902.
Copy the exact model ID from GPTProto Quick Start; do not assume the page title is the request string.
Re-test function schemas, structured JSON, streaming, and error recovery.
Check image and video input formatting if the workflow is multimodal.
Measure first-pass completion, retries, output tokens, end-to-end latency, and human corrections.
Keep the previous route available until the new revision passes regression tests.
A lower token rate is not automatically more affordable when retries or manual repair increase. Measure cost per accepted result.
Choose Qwen3.8-Max-0902 for large repositories, multimodal evidence, long outputs, structured tool use, or sustained work across many steps. It fits coding agents, research assistants, and enterprise systems expected to produce a verifiable deliverable.
Use a smaller or Flash-class model for short, latency-sensitive tasks. GPTProto's shared key and balance let developers compare the model catalog without separate provider accounts. Check live pricing and evaluation results before setting a default.
Guides, comparisons, and updates related to this model.
All Articles
Claude Fable 5.1 and Mythos 5.1 share one model but differ in access. See pricing, features, benchmarks, safeguards, and API migration changes.

Meta Description: Compare GLM 5.3 Flash vs DeepSeek V4 Flash for coding, frontend work, agents, speed, context, and API pricing to choose the better model.

Qwen3.8-Flash-Next vs GLM-5.3 Flash compared for coding, agents, frontend work, speed, pricing, context, licenses, and production use.

Compare 7 best AI gateways for developers in 2026 by pricing, routing, cost controls, deployment, and production trade-offs.