料金プラン+7% ボーナス
Michael Johnson2026-07-27

Qwen 3.8 Max vs Qwen 3.7 Max:何が変わり、開発者はどちらを選ぶべきか?

コーディング、コンテキスト、料金、APIの安定性、本番利用におけるQwen 3.8 MaxとQwen 3.7 Maxを比較します。開発者が3.8をテストしつつ、3.7を本番導入すべき理由を解説します。

Qwen 3.8 Max vs Qwen 3.7 Max:何が変わり、開発者はどちらを選ぶべきか?

2026年8月7日更新: Qwen3.8-Maxは、変更され続けるプレビューではなく、現在は安定した本番APIです。この比較における推奨事項も、それに合わせて更新しました。

Qwen3.8-Maxは、複雑性の高い新規ワークロードにおける、より強力なデフォルト選択肢になりました。ネイティブな画像・動画理解、95Bのアクティブパラメータを持つ2.4T MoEアーキテクチャ、長期タスクへの対応強化、最大128Kの出力を備えています。

Qwen3.7-Maxには、依然として魅力的な強みが1つあります。それは価格です。GPTProtoでは現在、入力100万トークンあたり0.36ドル、出力100万トークンあたり1.44ドルです。そのため、テキストのみのコーディング、文書分析、新モデルのマルチモーダル機能やエージェント機能の改善を必要としない既存の本番ワークフローに適しています。

結論:新しい複雑なコーディング、ビジュアル作業、長時間稼働するエージェントにはQwen3.8-Maxを選びましょう。最大性能よりも、テキスト処理のコスト効率や、すでに検証済みの本番ベースラインが重要な場合はQwen3.7-Maxを使い続けてください。

目次

Qwen 3.8 Max vs Qwen 3.7 Max at a Glance

Category Qwen 3.8 Max Qwen 3.7 Max
Status Stable production API Stable production API
Model name qwen3.8-max qwen3.7-max
Release date August 3, 2026 May 19, 2026
Architecture 2.4T MoE, 95B active Not publicly disclosed
Context window Up to 1M tokens Up to 1M tokens
Maximum output Up to 128K tokens Up to 128K tokens in current official listings
Inputs Text, images, and video Text
Thinking Hybrid thinking Thinking supported
Function calling Supported Supported
Official list price $2 input / $6 output per 1M $2.50 / $7.50; temporary $1.25 / $3.75 promotion
GPT Proto access Available now Available now
Best fit Complex coding, visual work, research, long agents Lower-cost text-only production
Open weights Announced; not released as of August 7 Proprietary

The 1M context figure for 3.8 matters because several early explainers still list it as unknown. QwenCloud’s current text-generation model table now lists 1M for both models. That closes one spec gap, but it does not show that 3.8 uses long context more accurately.

What Did Qwen 3.8 Max Actually Upgrade?

In its July 19 announcement, Qwen confirmed the 2.4T parameter count and said open weights would follow. Alibaba’s Qoder release note positions 3.8 as an improvement over 3.7 in coding and professional productivity, especially full-stack development, data analysis, office work, and other long-horizon tasks. A later Qwen update says the preview made a notable step on web frontend work.

Those are useful signals, but they are vendor statements rather than measured head-to-head results. No public table tells us how much better 3.8 is at resolving repository issues, completing tool-driven tasks, or preserving requirements across a long session. The active parameter count is also undisclosed, so the 2.4T headline does not tell developers the model’s serving cost or latency.

The honest upgrade story is therefore narrow: 3.8 targets better execution on the workloads 3.7 already targeted, while adding a larger model and a faster release cadence. It is not a context-window upgrade. It is not a function-calling upgrade. Whether it is a quality upgrade remains a reasonable hypothesis, not a verified result.

There is another cost. Qwen says the preview is improving continuously. That sounds attractive during evaluation, but a moving model complicates regression testing. The same prompt can behave differently after an unannounced update, and a passing test today does not guarantee the same result next week. For production teams, version stability is a feature.

Qwen 3.8 Max vs Qwen 3.7 Max for Coding

Qwen3.7-Max has a straightforward case for repository-scale coding. It accepts up to 1M tokens, supports thinking and tools, and is already callable through a conventional API. Its text-only interface is not a limitation when the agent receives source files, diffs, logs, and tool results as text.

Independent measurements also give us a baseline. Artificial Analysis scores Qwen3.7-Max at 46 on its Intelligence Index. It measured output speed at 202.2 tokens per second and time to first token at 2.62 seconds through Alibaba’s API. Those results will vary by provider and workload, but they are still more useful than a ranking claim without a methodology.

There is a trade-off. Qwen3.7-Max generated 100M output tokens during the Intelligence Index evaluation, compared with a 63M median for models in its comparison group. In plain language: it can be verbose. Fast token generation does not automatically mean a short response or a low total bill, especially when output tokens cost more than input tokens.

Qwen3.8-Max-Preview is more interesting for exploratory frontend and long-running agent work. Qwen specifically called out broad gains in web frontend, and Qoder positions it for full-stack development. I would test it on UI generation, multi-stage debugging, and tasks that mix code with office or data work. I would not move a production coding agent solely because of the model name.

A launch-week r/ClaudeCode discussion drew more than 40 comments. That is a demand signal, not a performance result. Community excitement can tell us which workloads developers care about; without controlled prompts and published outputs, it cannot tell us which model should carry production traffic.

For a fair Qwen 3.8 Max vs Qwen 3.7 Max coding test, keep the repository commit, system prompt, tools, permissions, and success criteria identical. Record tool failures, tests passed, human corrections, wall-clock time, and token use. If any of those inputs differ, the result is a demo—not a comparison.

Performance: Measured Results vs Vendor Claims

Alibaba describes Qwen3.8-Max-Preview as one of the leading frontier models and says it ranks behind only Fable 5 in its internal evaluation. The first half is plausible. The second half is not independently verifiable because Alibaba has not released the benchmark names, scores, prompts, or evaluation procedure behind it.

Evidence Qwen 3.8 Max Qwen 3.7 Max
Vendor positioning Improved coding and Cowork; internal ranking behind Fable 5 Agent-focused flagship for coding, productivity, and long autonomous work
Public benchmark table Not published Available vendor results, plus independent composite testing
Artificial Analysis Intelligence Index No result found as of July 27, 2026 46
Independently measured output speed Not available 202.2 tokens/s through Alibaba’s API
Independently measured TTFT Not available 2.62 seconds through Alibaba’s API
Reproducibility Preview changes during evaluation period More stable released endpoint

This evidence gap does not prove that 3.8 is worse. It means “which is better” has two answers. On expected capability, 3.8 is the likely winner. On demonstrated, reproducible performance, 3.7 currently has the stronger case.

Qwen 3.8 Max vs Qwen 3.7 Max Pricing

Qwen3.8-Max now uses standard pay-as-you-go token pricing. The earlier Preview-era Credits plan is no longer the relevant basis for this comparison.

Pricing Route Qwen 3.8 Max Qwen 3.7 Max
Official input price $2 per 1M tokens $2.50 per 1M tokens
Official output price $6 per 1M tokens $7.50 per 1M tokens
GPT Proto input price $1.80 per 1M tokens $0.36 per 1M tokens
GPT Proto output price $5.40 per 1M tokens $1.44 per 1M tokens

Although Qwen3.8-Max has a lower official list price than Qwen3.7-Max, GPT Proto currently offers a much deeper discount on the older model. On GPT Proto, Qwen3.7-Max costs 80% less for input and about 73% less for output than Qwen3.8-Max.

For example, a workload using 10 million input tokens and 2 million output tokens would cost:

  • Qwen3.8-Max on GPT Proto: 10 × $1.80 + 2 × $5.40 = $28.80

  • Qwen3.7-Max on GPT Proto: 10 × $0.36 + 2 × $1.44 = $6.48

That makes Qwen3.7-Max the clear budget choice for high-volume, text-only workloads. However, Qwen3.8-Max adds native image and video understanding, structured outputs, a longer output ceiling, and stronger positioning for complex coding and long-horizon agents.

The cheapest request is not always the cheapest completed task. Teams should compare first-pass success, retries, tool failures, token usage, and human correction time. Start with Qwen3.8-Max when the new capabilities affect the workflow; keep Qwen3.7-Max when its lower price delivers the better cost per accepted result.

Why There Is No Matched Output Table Here

A side-by-side screenshot would look convincing. It would also be misleading right now.

Qwen3.8-Max-Preview is not available on GPT Proto, and its official documentation says the model can change throughout the preview. Comparing a fresh 3.7 API run with an undated 3.8 screenshot from a different product would mix model versions, infrastructure, tools, and possibly system prompts. I will not label that as a controlled test.

The publication-quality test is simple: run the same repository task against both models on the same day, preserve the full prompt and tool trace, and publish the raw outputs alongside the judgment criteria. Until that run exists, the absence of a sample is more honest than a decorative winner badge.

Which Model Should You Choose?

Your Priority Better Choice Why
Best overall Qwen capability Qwen 3.8 Max New flagship with stronger coding, visual, research, and long-horizon positioning
Complex repository or agent work Qwen 3.8 Max Designed for longer autonomous execution and closed-loop tool workflows
Image, video, or visual-document input Qwen 3.8 Max 3.7 Max is text-only
Lowest GPT Proto text-token cost Qwen 3.7 Max $0.36 input and $1.44 output per 1M tokens
Existing regression-tested application Keep 3.7 until validated A stable release still needs workload-specific migration testing
New application starting today Start with 3.8 There is no longer a preview-status reason to default to 3.7
Self-hosting today Neither 3.7 is proprietary; 3.8 weights are announced but not yet downloadable

The old recommendation—“ship 3.7 and only evaluate 3.8”—no longer fits the product state.

For new work, start the evaluation with Qwen3.8-Max on GPT Proto. Keep Qwen3.7-Max when its much lower token price or an existing validated baseline produces a better cost per accepted task.

Try Qwen 3.8 Max Through GPT Proto

Qwen3.8-Max is now available through GPT Proto’s OpenAI-compatible API.

export GPTPROTO_API_KEY="your_api_key_here"

curl https://gptproto.com/v1/chat/completions \
  -H "Authorization: Bearer $GPTPROTO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-max",
    "messages": [
      {
        "role": "system",
        "content": "Inspect the existing implementation before proposing changes. Preserve current API contracts and list the verification steps for every edit."
      },
      {
        "role": "user",
        "content": "Review this repository architecture and identify the three highest-risk assumptions."
      }
    ]
  }'

The Python version uses the official OpenAI client:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["GPTPROTO_API_KEY"],
    base_url="https://gptproto.com/v1",
)

response = client.chat.completions.create(
    model="qwen3.7-max",
    messages=[
        {
            "role": "system",
            "content": (
                "You are a senior software engineer. Return the smallest "
                "safe patch and explain each changed file."
            ),
        },
        {
            "role": "user",
            "content": (
                "Review this function for correctness and propose a tested fix: "
                "def divide_total(total, count): return total / count"
            ),
        },
    ],
)

print(response.choices[0].message.content)

The SDK adds the Bearer authorization header. Switching to another compatible model only requires changing the model string; keep a separate evaluation set so a convenient swap does not become an untested production change.

A Safe Upgrade Checklist

Start by saving a 3.7 baseline from your real workload, not a synthetic coding puzzle. Preserve the prompt, repository state, tools, expected output, latency, and token usage. Then run the same package against 3.8 Preview and inspect failures, not only the best-looking answer.

Set the migration threshold before seeing the result. For example, require the new model to pass the same tests with no increase in tool errors and no unacceptable cost change. Re-run the evaluation after major preview updates. Finally, wait for a stable model identifier and a published price before moving traffic that carries an uptime or budget commitment.

This process is less exciting than switching on launch day. It is also how you keep a model upgrade from becoming an incident.

Final Verdict

Qwen3.8-Max is now the better default Qwen model for new complex workloads. The stable API removes the largest objection in the original comparison, while multimodal input and stronger long-horizon behavior give it a broader capability ceiling.

Qwen3.7-Max is not obsolete. Its GPT Proto price makes it a strong text-only production model, especially when a workload is already tested and does not benefit from vision or extended autonomous execution.

Start new evaluations with Qwen3.8-Max. Keep 3.7 where it wins on cost per accepted task—not simply because it was released earlier.

1つのキーで、より多くのAIモデル

1つのOpenAI互換APIを通じて、主要AIモデルを手頃な価格で利用できます。

APIモデルを閲覧
1つのキーで、より多くのAIモデル
関連モデル
すべてのモデル
Qwen
by Qwen
10% OFF
Qwen
by Qwen
10% OFF
OpenAI
20% OFF
Claude
10% OFF

よくある質問

Qwen 3.7 MaxからQwen 3.8 Maxへのアップグレード内容は?

Qwenは、3.8について、コーディング、フルスタック開発、データ分析、オフィスワークフロー、長期タスクにおける改善を位置付けています。また、総パラメータ数は2.4Tです。両モデルともすでに100万トークンのコンテキスト、思考、関数呼び出し、組み込みツールを提供しているため、主張されているアップグレードは新しいAPI機能というより、主に実行品質の向上です。

コーディングではQwen 3.8 MaxのほうがQwen 3.7 Maxより優れていますか?

一部のワークロードではおそらく優れていますが、現時点でその結論を独立測定したデータはありません。Qwenはフロントエンドとコーディングの改善を報告していますが、現在、公開されている根拠がより強いのは3.7です。新規開発では3.8をテストし、プレビューが安定するまでは本番で3.7を使い続けてください。

Qwen 3.8 Maxのコンテキストウィンドウは100万トークンですか?

はい。QwenCloudの現在のテキスト生成モデル一覧では、Qwen3.8-Max-PreviewとQwen3.7-Maxの両方に100万トークンのコンテキストウィンドウが記載されています。

Qwen 3.8 Maxの料金はいくらですか?

トークン単位の単独価格はまだありません。Token PlanとQoderのクレジットベースのアクセスで利用できます。Token Planでは現在、個人向けの期間限定プランとして月額6ドル、18ドル、68ドルが掲載されていますが、これらのサブスクリプションは複数モデルを対象としており、3.7のトークン料金と直接比較することはできません。

Qwen 3.8 Maxはオープンソースですか?

現時点では違います。Qwenはオープンウェイトを近日公開すると述べていますが、ウェイト、リリース日、ライセンスはいずれも公開されていません。Qwen3.7-Maxはプロプライエタリです。

Qwen 3.8 MaxはGPTProtoで利用できますか?

はい。GPTProtoでは現在、Qwen3.8-Maxを利用できます。

開発者はどのモデルを選ぶべきですか?

安定した本番アクセス、独立した性能データ、予測可能なトークンコストを重視するならQwen3.7-Maxを選びましょう。フロントエンド、Cowork、長時間のエージェントタスクではQwen3.8-Max-Previewを評価し、Alibabaが安定版、ベンチマーク表、トークン単価を公開した時点で移行を再検討してください。
Qwen 3.8 Max vs GLM 5.2:2026年のコーディングにはどちらが優れている?

Qwen 3.8 Max vs GLM 5.2:2026年のコーディングにはどちらが優れている?

2026年8月7日更新: Qwen3.8-Maxは現在、安定した本番用APIとなっており、GPTProtoから利用できます。以前のPreview期におけるデプロイ推奨事項は更新されています。 要約 ホスト型モデルとしての最高性能、マルチモーダル入力、フロントエンド開発、画像分析、または長期的なエージェント実行を重視する場合は、Qwen3.8-Maxを選びましょう。 低いトークンコスト、MITライセンスの重み、自社ホスティング、または再現可能なオープンなデプロイを重視する場合は、GLM-5.2を選びましょう。 現在、両モデルとも安定したAPIアクセスを利用できます。GLMだけが本番利用の選択肢という状況ではありません。 Qwenの公式料金は入力100万トークンあたり$2、出力100万トークンあたり$6です。GPTProtoでは現在、GLM-5.2を入力100万トークンあたり$1.26、出力100万トークンあたり$3.96で掲載しています。 Qwenは性能を最優先する場合のより強力な選択肢です。GLMはコストと制御性を重視する場合のより強力な選択肢です。

Michael Johnson | 2026-07-22

Qwen 3.8 Max vs Kimi K3:実際のコーディング業務に対応できるのはどちらか?

Qwen 3.8 Max vs Kimi K3:実際のコーディング業務に対応できるのはどちらか?

更新 — 2026年7月28日:Moonshot AIは、Kimi K3の完全な重み、モデルカード、カスタムライセンス、技術レポートを公開しました。これにより、Kimi側の提供状況に関する疑問は解消されました。ただし、2.8兆パラメータモデルのセルフホスティングが容易になったわけではありません。公式リポジトリは約1.56 TBあり、Moonshotは64基以上のアクセラレーターを備えたスーパーノード構成を推奨しています。 Qwen 3.8 MaxとKimi K3の比較は、2つの巨大な中国製AIモデルによる単純な対決に見えます。Alibabaの2.4兆パラメータのプレビュー版と、Moonshot AIの2.8兆パラメータのフラッグシップモデルです。数字からは単純な結論が導かれます。より大きなモデルが勝つはずだ、と。 しかし、入手可能な証拠はそのような結果を示していません。また、開発者にとって最も有用な比較でもありません。 2026年7月23日時点で、Qwen 3.8 MaxはAlibabaのToken Planを通じて提供される、なお変化の続くプレビューです。一方、Kimi K3にはすでに文書化されたAPI、公開済みのトークン価格、100万トークンのコンテキストウィンドウ、完全な重みをリリースする日付付きの計画があります。性能差は小さいかもしれません。しかし、製品としての準備状況には差があります。 私の判断は明快です。今日、実際のアプリケーションを構築し、予算を立てる必要があるなら、Kimi K3のほうが安全な選択です。Qwen 3.8 Max Previewは、特にAlibabaのプロモーションCreditsによって低コストで試せる間は、コーディングワークフローでテストする価値があります。しかし、本番採用を決めるには、まだ安定した情報が十分に提供されていません。 要約:現時点で本番環境に安全なのはKimi K3 従来型のAPI、予測可能なトークン単価、ネイティブな画像・動画理解、または今すぐ顧客向け製品の背後に配置できるモデルが必要なら、Kimi K3を選びましょう。すでにAlibabaのコーディングエコシステムを利用しており、将来性のある新モデルを低いプロモーションコストで試したいなら、Qwen 3.8 Max Previewを選びましょう。 公開時点で利用できた唯一の詳細な同条件のコーディングテストでは、Kimi K3が83点、Qwen 3.8 Maxが80点でした。この3点差は有用な証拠ですが、普遍的な順位を示すものではありません。テストでは、Qwenはより明確なシステム境界と完璧なツール実行を示し、Kimiは修正履歴と再生成をより完全に処理しました。両モデルとも、事実確認による修正が必要な、裏付けのない推論を行いました。 平たく言えば、現時点で導入判断ではKimiが勝っています。Qwenは性能競争に敗れたわけではありません。勝ったと宣言するには、まだ早すぎるのです。

Schuyler Stacy | 2026-07-28

2026年開発者向け最適なAI API:10プラットフォーム比較

2026年開発者向け最適なAI API:10プラットフォーム比較

TL;DR 最適な直接API: OpenAIは汎用用途における最も安全なデフォルトです。Anthropic Claudeはコーディングと長時間稼働するエージェントに最適です。Geminiは低コストのマルチモーダルプロトタイピングに適しており、DeepSeekはテキストトークン単価で優位に立っています。 最適なマルチモデルオプション: 多数のLLMをテストするなら、OpenRouterが最も明確な選択肢です。1つの製品でテキスト、画像、動画モデルを1つのAPIキーと共通残高で利用する場合は、GPTProtoの方が適しています。 最適なインフラストラクチャ: Amazon BedrockはAWSのガバナンス下にあるエンタープライズ環境に適しています。一方、Replicate、fal.ai、Together AIはオープンモデルや生成メディアの推論に適しています。 すべての用途における唯一の勝者は存在しません。ワークロードへの適合性、モデルのカバレッジ、実際の課金単位、本番運用の制御機能、移行コストを比較してください。価格と提供状況は2026年7月14日時点で確認しています。導入前に各プロバイダーの最新ページを確認してください。

Tiffany Layne | 2026-07-15

Qwen 3.8 Maxとは?リリース、仕様、料金、オープンウェイト

Qwen 3.8 Maxとは?リリース、仕様、料金、オープンウェイト

2026年8月7日更新: Alibabaは8月3日、Qwen3.8-Maxの本番版を正式リリースしました。これにより、従来の qwen3.8-max-preview に代わり、現在のフラッグシップAPIモデルとなっています。本記事では、確定したアーキテクチャ、コンテキストウィンドウ、料金、APIの提供状況、オープンウェイトのスケジュールを反映しています。 Qwen3.8-Maxは、これまでで最も高性能なQwenモデルです。総パラメータ数2.4兆、リクエストごとのアクティブパラメータ数950億を備えた、マルチモーダルSparse Mixture-of-Expertsモデルです。 安定版モデルは、テキスト、画像、動画の入力に対応し、テキストを出力します。最大100万トークンのコンテキストウィンドウと、最大128Kトークンの出力に対応しています。複雑なコーディング、視覚分析、研究、専門業務、長時間にわたるエージェントタスク向けに設計されています。 今回のリリースによって、開発者にとっての実際の選択肢も変わりました。Qwen3.8-Maxは、仕様が変わり続けるプレビュー版やCredits限定の個人プランに制限されなくなりました。通常の従量課金APIが利用でき、Alibabaは入力100万トークンあたり2ドル、出力100万トークンあたり6ドルと案内しています。また、 GPTProtoのQwen 3.8 Max API からも利用できます。 私の簡単な見解:Qwen 3.8 Maxは、完成した製品リリースではなく、実在する非常に興味深いプレビューです。開発者はテストを行い、すべての結果の日付を記録し、Alibabaがまだ公開していない仕様を前提に本番移行を計画することは避けるべきです。

Tiffany Layne | 2026-07-23