Schuyler Stacy2026-07-28

Qwen 3.8 Max vs Kimi K3:本格的なコーディング作業に対応できるのはどちらか?

Qwen 3.8 MaxとKimi K3を、コーディング、API料金、コンテキスト、マルチモーダル対応、オープンウェイトの観点で比較し、どちらのモデルが導入準備を整えているかを確認します。

Qwen 3.8 Max vs Kimi K3:本格的なコーディング作業に対応できるのはどちらか?

更新 — 2026年7月28日:Moonshot AIは、Kimi K3の完全な重み、モデルカード、独自ライセンス、技術レポートを公開しました。これにより、Kimi側の提供状況に関する疑問は解消されました。ただし、2.8兆パラメーターのモデルをセルフホストしやすくなったわけではありません。公式リポジトリの容量は約1.56 TBで、Moonshotは64基以上のアクセラレーターを備えたスーパーノード構成を推奨しています。

 

Qwen 3.8 MaxとKimi K3の比較は、中国の巨大AIモデル2つによる分かりやすい競争に見えます。Alibabaの2.4兆パラメーターのプレビュー版と、Moonshot AIの2.8兆パラメーターのフラッグシップモデルです。この数字を見ると、単純な結論に誘導されます。より大きいモデルが勝つはずだ、と。

しかし、現在得られる証拠はそのような結論を示しておらず、開発者にとって最も有用な比較でもありません。

2026年7月23日時点で、Qwen 3.8 MaxはAlibabaのToken Planを通じて提供される、なお変化し続けるプレビューです。一方、Kimi K3には文書化されたAPI、公開されたトークン価格、100万トークンのコンテキストウィンドウ、完全な重みをリリースする日程があります。能力差は小さいかもしれません。しかし、製品としての準備度には明確な差があります。

私の判断は明快です。今日、実際のアプリケーションを構築し、予算を立てる必要があるなら、より安全な選択はKimi K3です。Qwen 3.8 Max Previewは、特にAlibabaのプロモーションCreditsによって低コストで試せる間は、コーディングワークフローで検証する価値があります。しかし、本番環境での採用を決めるには、まだ安定した情報が十分ではありません。

要約:現時点で本番利用に安全なのはKimi K3

一般的なAPI、予測可能なトークン単価、画像と動画のネイティブ理解、または今すぐ顧客向け製品の背後に配置できるモデルが必要なら、Kimi K3を選びましょう。Alibabaのコーディングエコシステムをすでに利用しており、将来性のある新モデルを低いプロモーション価格で試したいなら、Qwen 3.8 Max Previewを選びましょう。

公開時点で利用できた唯一の詳細な同条件コーディングテストでは、Kimi K3が83点、Qwen 3.8 Maxが80点でした。この3点差は有用な証拠ですが、普遍的な順位ではありません。テストではQwenがより明確なシステム境界と完璧なツール実行を示し、Kimiは改訂履歴と再生成をより完全に処理しました。両モデルとも、事実確認による修正が必要な根拠のない推論を行いました。

簡単に言えば、現時点の導入判断ではKimiが勝っています。Qwenが能力競争に敗れたわけではありません。勝利を宣言するにはまだ早いのです。

目次

Qwen 3.8 Max vs Kimi K3 at a Glance

Category Qwen 3.8 Max Kimi K3
Current status qwen3.8-max-preview; may change during Preview or be replaced by a production version Released model with documented API access
Reported total parameters 2.4T 2.8T
Architecture details Active parameters and full MoE configuration not yet disclosed KDA, Attention Residuals, and Stable LatentMoE; 16 of 896 experts activated
Context window Not clearly disclosed on the current model product page 1,048,576 tokens
Documented inputs Text and images Text, images, and video
Reasoning Always on; low, high, or xhigh; default xhigh Always on; low, high, or max; default max
Current access model Alibaba Token Plan and supported coding or agent tools Conventional hosted API, Kimi products, and coding tools
Ordinary per-token price Not published $3/M fresh input and $15/M output at Moonshot’s current list rate
GPT Proto price Not yet available $2.70/M input and $13.50/M output
Open-weight status on July 28 Announced as coming “soon”; no public checkpoint, date, or license yet Full weights released in `moonshotai/Kimi-K3` under the custom Kimi K3 License
Self-hosting reality Cannot be assessed until Alibaba publishes the checkpoint and deployment guidance Technically available; the 1.56 TB repository and 64+ accelerator recommendation make it a data-center deployment, not a laptop model
Best current fit Low-cost Preview testing in supported coding workflows Production applications, multimodal agents, and costed coding workloads
Main risk Preview behavior, access terms, and commercial pricing may change Always-on reasoning can increase latency, output tokens, and long-session cost

Alibaba’s own Token Plan documentation says the Qwen Preview may be continually updated and later removed or replaced. Moonshot’s Kimi K3 announcement and API guide disclose substantially more about the model and its current interface.

Why This Is Not an Equal Comparison Yet

Qwen 3.8 Max Is Still a Moving Preview

Alibaba describes Qwen 3.8 Max as a 2.4T model, but the hosted model available today is specifically qwen3.8-max-preview. That suffix matters. Alibaba states that its capabilities may be upgraded throughout the Preview and that the endpoint may later be removed or replaced by a production version.

This makes every current benchmark a dated snapshot. A useful Qwen result should record the test date, endpoint, reasoning level, agent shell, tool permissions, and raw output. Without those details, two people may believe they tested the same model while actually receiving different Preview behavior.

The access terms also limit what developers can conclude from the low launch price. Alibaba’s Personal Token Plan currently offers Qwen 3.8 Max through a Credits-based subscription, but its official terms restrict the key to supported interactive coding and agent tools. Automation scripts, custom application backends, and non-interactive batch calls are prohibited on that plan.

So Qwen is callable. It is not yet a normal pay-as-you-go backend for a customer-facing app.

Kimi K3 Has Completed Its Open-Weight Release

Kimi K3 is now available in two forms: a conventional hosted API and a downloadable checkpoint. Moonshot AI has published the full weights in the official `moonshotai/Kimi-K3` repository, together with a model card, deployment guidance, a custom license, and the K3 technical report.

The release adds details that were unavailable when this comparison was first published. K3 has 104B activated parameters per token, uses MXFP4 weights with MXFP8 activations, and is currently documented for vLLM, SGLang, and TokenSpeed. The repository is about 1.56 TB across 96 safetensors shards.
That is a real open-weight release. It is not the same as an unrestricted MIT release or an easy local install. The Kimi K3 License adds conditions for large Model-as-a-Service businesses and very large commercial products, while Moonshot recommends supernode configurations with 64 or more accelerators for deployment.
In plain English: Kimi has removed the ownership uncertainty. It has not removed the infrastructure cost.

Coding Comparison: What the 269-File Test Actually Shows

The best current head-to-head evidence is a matched StackPerf architecture test. Both models received the same read-only snapshot of two unfamiliar software projects containing 269 files. They used OpenCode 1.17.13, identical task instructions, the same permissions, a 60-minute limit, and a 65,536-token completion ceiling. Editing, web access, plugins, and subagents were disabled.

The final reports were anonymized before scoring, then checked separately for unsupported claims.

Test result Qwen 3.8 Max Kimi K3
Final score 80/100 83/100
Tool calls 44 53
Failed or denied tool calls 0 2, followed by successful recovery
Repository citation occurrences 354 274
Unsupported claim groups after verification 7 7
Clearest strength System boundaries and replay metadata Revision, regeneration, and lifecycle state

Where Kimi K3 Did Better

Kimi’s report modeled the editing lifecycle more completely. It tracked revisions, superseded takes, retry attempts, fallback history, artifact hashes, invalidation, editorial selection, and trimming. That made its proposed contract better suited to a system in which one scene may be generated, rejected, revised, and generated again without losing its history.

Kimi also finished sooner and used fewer tokens on the tested route. That result is encouraging for long coding tasks, but it includes the serving system around the model. Kimi ran through a Kimi subscription endpoint, while Qwen ran through Alibaba’s international Token Plan endpoint. Provider capacity, caching, and launch-period traffic all affected the observed result.

My reading is that Kimi won this task because it understood changing state more completely, not because it wrote prettier code or proved itself better at every software task.

Where Qwen 3.8 Max Did Better

Qwen drew a cleaner boundary between the two systems and created a stronger replay record. Its proposed metadata included the model, seed, prompt, references, provider request ID, media location, duration, and cost. Those fields matter when a team needs to reproduce an output or investigate why the same workflow changed after a provider update.

Its tool behavior was also disciplined: 44 tool calls, all successful. Qwen produced the longer report with fewer requests and received a higher tool-use score in the blind review.

The cost of that cleaner architecture was weaker lifecycle modeling. Its contract did not fully cover scene order, revisions, supersession, retry history, or take invalidation. Clean boundaries are valuable. They are not enough when the product’s state keeps changing.

What This Test Cannot Prove

One architecture-analysis run cannot tell us which model produces better frontend code, fixes more real bugs, passes more tests, or needs fewer human corrections over a month of development. It also cannot isolate model intelligence from endpoint performance.

The most important warning is hidden in the factual review: both models cited real repository locations, yet both produced seven groups of claims that went beyond what the code proved. A report can contain hundreds of valid citations and still reach an unsafe conclusion.

That is the real takeaway. Use either model as an engineering assistant. Do not treat either one as its own reviewer.

How Developers Should Test Qwen 3.8 Max Against Kimi K3

Test Give Both Models Measure
Repository architecture review The same frozen commit, architecture question, read-only tools, and time limit Correct file citations, missed dependencies, unsupported claims, and review time
Multi-file implementation The same issue, tests, writable files, and tool permissions Tests passed, files changed, retries, regressions, and human corrections
Visual frontend repair The same screenshot, source files, browser tools, and target behavior Visual match, valid code, repair loops, and final test result

Keep the agent shell and permissions identical. Set a fixed time limit. Record the full model ID and date, especially for Qwen’s moving Preview. Then capture task completion, wall-clock time, input and output tokens, cache hits, failed tool calls, retries, human interventions, and final tests passed.

Do not score an answer because it “looks thorough.” Check whether the patch works and whether the model’s claims survive review.

For high-value architecture decisions, there is another useful pattern: run both models independently, hide their identities during review, and compare their disagreements. The 269-file test produced its strongest design only after combining Qwen’s system boundaries with Kimi’s lifecycle model. Sometimes the right answer to Qwen versus Kimi is both, followed by verification.

Qwen 3.8 Max vs Kimi K3 Pricing and Cost

Kimi K3 Has Predictable Token Pricing

Kimi K3 has an ordinary token-based price. The current GPT Proto Kimi K3 API page lists $2.70 per 1M input tokens and $13.50 per 1M output tokens, 10% below Moonshot’s current $3/$15 list rate.

That makes a basic cost estimate possible before a test begins:

estimated cost = (input tokens ÷ 1,000,000 × $2.70) + (output tokens ÷ 1,000,000 × $13.50)

A run with 200,000 input tokens and 20,000 output tokens would cost about $0.81 at the current GPT Proto rate: $0.54 for input and $0.27 for output. Retries, additional turns, tool results, and retained reasoning history would increase the total.

Predictable does not mean automatically cheap. Kimi always reasons, and reasoning content counts toward token use. Long agent histories can therefore become expensive even when the visible final answer is short.

Qwen 3.8 Max Is Cheaper to Experiment With, but Harder to Budget

Qwen 3.8 Max Preview currently uses subscriptions and Credits rather than a disclosed per-million-token price. Alibaba’s Personal Token Plan lists promotional monthly prices of CNY 39, CNY 139, and CNY 499 for its three tiers. During the Preview, Qwen calls may consume as little as 10% of the standard Credits rate, with an additional time-limited night discount documented for calls between 22:00 and 08:00 UTC+8.

That is attractive for interactive testing. It is not a clean cost comparison.

Credits can vary with token use, caching, reasoning, and tools, so a CNY subscription cannot honestly be converted into a stable $X per 1M tokens figure from the information currently published. The Personal Plan restriction also means that its promotional economics do not represent the cost of running Qwen behind your own application.

Qwen wins on low-cost Preview access. Kimi wins on production cost visibility.

Open Weights and Self-Hosting

Kimi K3 now has a concrete advantage over Qwen 3.8 Max: its exact public checkpoint can be downloaded today. Moonshot AI has released the full 2.8T-parameter weights, model card, license, and technical report. Alibaba has announced open weights for Qwen 3.8 Max, but the exact hosted Max checkpoint, release date, license, and deployment requirements remain unconfirmed.
The word “open” still needs qualification. Kimi K3 is open-weight under the custom Kimi K3 License, not MIT. The license broadly permits use, modification, deployment, fine-tuning, and redistribution, but adds a separate-agreement requirement for Model-as-a-Service businesses above the stated revenue threshold and an attribution requirement for very large commercial products.
Hardware is the other boundary. The official K3 repository is about 1.56 TB, and Moonshot recommends supernode deployments with at least 64 accelerators. That makes self-hosting plausible for well-funded infrastructure teams, not a substitute for an API key on a developer workstation.
Choose the hosted Kimi K3 API when you want predictable token billing and no cluster operations. Choose the weights when data control, custom inference, fine-tuning, or deployment ownership is worth operating the hardware and reviewing the license. For Qwen 3.8 Max, wait until Alibaba publishes the actual checkpoint before making the same calculation.

Multimodal Inputs, Reasoning, and Agent Behavior

Capability Qwen 3.8 Max Preview Kimi K3
Image understanding Documented Documented
Video understanding Not clearly documented as a current model input Documented
Thinking mode Always on Always on
Reasoning levels low, high, xhigh low, high, max
Default reasoning level xhigh max
Context window Not clearly disclosed on the current product page 1,048,576 tokens
Structured output Preview capabilities require continued verification JSON mode and strict JSON Schema documented
Long tool histories Behavior may change with the Preview Full assistant messages, including reasoning content, should be preserved

Qwen’s documented vision support makes it relevant for screenshot-based debugging. Kimi goes further by accepting video, which is useful when the input is a screen recording, animation reference, or product demo that would be difficult to describe frame by frame.

Kimi’s richer documented interface also adds integration work. Its Preserved Thinking behavior means a multi-turn application should return the complete assistant message, including reasoning content and tool calls, rather than keeping only the visible answer. That history occupies context and is billed. The 1M-token window is large, but it is not free storage.

Both models default to their highest reasoning setting. For evaluation, keep those settings consistent. For production, test lower effort on simpler tasks. A model that solves 99% of requests at lower effort may be cheaper and faster than one left at maximum reasoning for every autocomplete, classification, or short transformation.

Which Model Should You Choose?

Your situation Better current choice Why
Building a customer-facing application now Kimi K3 Conventional API and forecastable token price
Running a long repository task with image or video input Kimi K3 1M context and documented image/video support
Testing inside Qwen Code, Qoder, or another supported Alibaba tool Qwen 3.8 Max Preview Low promotional Credits consumption
Need stable repeatable benchmarks Kimi K3, for now Qwen’s Preview may change between runs
Need the cleanest architecture and replay metadata Test Qwen It showed a real advantage in the matched architecture review
Need strong revision and regeneration handling Test Kimi It was more complete in the available matched test
Need downloadable weights today Kimi K3 The full checkpoint is public; Qwen 3.8 Max still has no released weights
Need to compare several model families behind one account Kimi K3 on GPT Proto One key and balance can access the broader model catalog

For an application going live this week, I would choose Kimi K3. It has enough published information to estimate cost, define integration behavior, and repeat a test against a stable model name.

For an internal coding experiment, I would not ignore Qwen. Its Preview handled a long repository analysis with zero failed tool calls and showed better architectural boundaries than Kimi in the matched test. That is a serious capability signal. It is not yet a production contract.

How to Try Kimi K3 Through GPT Proto

Qwen 3.8 Max is not yet available on GPT Proto, so the current integration example uses Kimi K3 only. You can call it through GPT Proto’s OpenAI-compatible Chat Completions endpoint with the kimi-k3 model string:

curl https://gptproto.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: $GPTPROTO_API_KEY" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Review this migration plan. Identify unsupported assumptions, missing rollback steps, and the tests required before production."
      }
    ]
  }'

The endpoint and authorization structure follow the GPT Proto API quickstart. For multi-turn Kimi workflows, preserve the complete assistant message returned by the API, including reasoning and tool-call fields, rather than storing only the final visible text.

You can try the Kimi K3 API, browse 200+ AI models, or use the GPT Proto unified AI API to compare Kimi with other text, image, video, and audio models. When Qwen 3.8 Max is added, the useful test will be the same task, prompt, tool permissions, and scoring method on both model endpoints.

Final Verdict

Qwen 3.8 Max and Kimi K3 appear close enough in coding capability that one small benchmark should not decide the next year of your architecture. Qwen’s cleaner system boundaries and flawless tool execution deserve attention. Kimi’s stronger lifecycle reasoning and three-point lead deserve attention too.

The product decision is now clearer. Kimi K3 combines a standard API, public token pricing, a documented 1M context window, multimodal input, and a released open-weight checkpoint. Qwen 3.8 Max remains a changeable Preview with Credits-based access, no ordinary production price, and no public weights.
Kimi K3 is the better model to deploy today, whether you begin with its hosted API or have the infrastructure to manage the checkpoint yourself. Qwen 3.8 Max may still become the stronger model, but developers should wait for its production endpoint, ordinary pricing, model card, license, and actual weight release before making that bet.

クリエイティブスタジオ

本番環境向けAPIを使用して、画像や動画などを生成します。

作成を開始する
クリエイティブスタジオ
関連モデル
すべてのモデル
MoonshotAI
10% OFF
Claude
20% OFF
Google
40% OFF
Google
40% OFF

よくある質問

コーディングではQwen 3.8 MaxとKimi K3のどちらが優れていますか?

Qwen 3.8 Maxが一般的に優れていると言えるだけの証拠はありません。利用可能な269ファイルのアーキテクチャテストでは、Kimi K3が83点、Qwenが80点でした。Qwenはより明確なシステム境界を作り、44回中44回のツール呼び出しに成功しました。一方、Kimiは改訂と再生成をより完全に処理しました。選択する前に、自分の実装タスクと修正タスクで両方をテストしてください。

Qwen 3.8 MaxとKimi K3ではどちらが安いですか?

AlibabaのToken Planによるプロモーション目的の実験では、現在Qwenのほうが安価です。ただし、そのCreditsは公開された100万トークン単価に換算できません。Kimiは本番環境の予算を立てやすいモデルです。GPTProtoでは現在、Kimi K3を入力100万トークンあたり$2.70、出力100万トークンあたり$13.50で掲載しています。

Qwen 3.8 MaxとKimi K3はオープンソースですか?

Kimi K3は独自のKimi K3 Licenseの下で、現在オープンウェイトとして提供されています。このライセンスは幅広い利用と変更を許可しますが、大規模なModel-as-a-Service事業や非常に大規模な商用製品に条件があるため、MITライセンスとは表現すべきではありません。Qwen 3.8 Maxはオープンウェイトを発表していますが、チェックポイントやライセンスはまだ公開していません。

Qwen 3.8 Maxは本番APIで利用できますか?

現在のQwen 3.8 Max PreviewはAlibabaのToken Planで利用できますが、Personal Planの規約では、そのキーを自動化スクリプト、カスタムアプリケーションのバックエンド、非対話型バッチ呼び出しに使用することを禁止しています。顧客向けバックエンドを設計する前に、文書化された本番利用経路と通常の商用価格が提供されるのを待ってください。

長時間のコーディングタスクにはどちらのモデルが適していますか?

Kimi K3は、文書化された100万トークンのコンテキストウィンドウ、一般的なAPI、利用可能な同条件テストでより少ないトークンでの結果を備えているため、現時点でより安全な選択です。Qwenもテストする価値があります。同じ60分間のリポジトリタスクを、失敗したツール呼び出しなしで完了しました。長いコンテキスト容量だけでタスク完了が保証されるわけではないため、合格したテスト、再試行、人手による修正を測定してください。

Kimi K3は画像や動画を理解できますか?

はい。Moonshotは、Kimi K3がテキスト、画像、動画をネイティブに理解すると説明しています。モデルはテキストを返すため、スクリーンショットベースのデバッグ、動画分析、インターフェースレビュー、その他のマルチモーダルなコーディングや知識業務に適しています。

Qwen 3.8 MaxまたはKimi K3はローカルで実行できますか?

Kimi K3の重みはダウンロードできますが、実用的なセルフホスティングにはデータセンター級のインフラが必要です。公式リポジトリは約1.56 TBで、Moonshotは64基以上のアクセラレーターを推奨しています。Qwen 3.8 Maxは、Alibabaが重みを公開するまでセルフホスティングできません。
Kimi K3とは?GPT-5.6やFable 5に本当に近いのか?

Kimi K3とは?GPT-5.6やFable 5に本当に近いのか?

TL;DR Kimi K3は、長時間にわたるコーディング、ナレッジワーク、推論、エージェントワークフロー向けにMoonshot AIが開発した、2.8兆パラメータのマルチモーダルモデルです。独立したテストでは、総合的にClaude Opus 4.8やGPT-5.5に近い位置にありますが、GPT-5.6 SolとClaude Fable 5には依然として及びません。K3はエージェントベンチマークで差を縮め、一部の自動化テストではトップに立っていますが、測定されたハルシネーション率はK2.6から上昇しています。 Kimi K3は現在、オープンウェイトとして公開されています。Moonshot AIは完全なチェックポイント、モデルカード、技術レポート、独自のKimi K3 Licenseを公開しました。公式Hugging Faceリポジトリは96個のsafetensorsシャードで約1.56 TBあり、Moonshotは64基以上のアクセラレータを備えたスーパーノード構成を推奨しています。オープンウェイト化によって、所有権に関する疑問は解消されました。ただし、K3が一般的なローカルモデルになったわけではありません。 ほとんどの開発者にとって、ホステッドAPIが現実的な出発点です。 GPTProtoのKimi K3 API は現在、入力トークン100万個あたり2.70ドル、出力トークン100万個あたり13.50ドルと表示されています。データ管理、カスタム推論、モデル変更に、インフラ構築やライセンス確認のコストをかける価値がある場合はウェイトを選びましょう。 要するに、Kimi K3はGPT-5.6やFable 5と同じ土俵で語れるほど近い性能を持ち、オープンウェイトでのリリースによって、どちらのクローズドモデルにもないデプロイメントの選択肢を開発者に提供しています。

Michael Johnson | 2026-07-28

コーディングにおけるGLM-5.2 vs Kimi K3:2026年、開発者にとって優れているのはどちら?

コーディングにおけるGLM-5.2 vs Kimi K3:2026年、開発者にとって優れているのはどちら?

TL;DR: 難しいタスク、長時間にわたるタスク、またはビジュアル要素を含むタスクでは、Kimi K3のほうが優れたコーディングモデルです。Moonshotが公開したコーディング比較ではGLM-5.2を上回り、ホスト型サービスを通じて画像と動画も扱えます。一方、日常的なリポジトリ作業のデフォルトとしては、GLM-5.2のほうが適しています。コストが大幅に安く、運用するモデルも小さく、寛容なMITライセンスを採用しているためです。Kimi K3も現在は重みが公開されていますが、1.56 TBのリポジトリ、64基以上のアクセラレータを推奨する構成、独自ライセンスにより、セルフホスティングへの取り組みは大幅に大きくなります。ボトルネックが能力ならKimiを、毎日のコストと運用の簡便さを重視するならGLMを選びましょう。 GLM-5.2とKimi K3 Codeの比較で興味深いのは、どちらもReactコンポーネントを作成したり、短いアルゴリズムを解いたりできることではありません。このレベルのモデルなら、その基準はすでにクリアしています。重要なのは、課題が複雑になったときにどうなるかです。リポジトリの監査、複数ファイルにまたがる移行、スクリーンショットでしか現れないバグ、あるいは複数のシステムの整合性を保つ必要がある、プレイ可能なThree.jsプロトタイプなどです。 価格差が重要になり始めるのも、まさにこの領域です。最も難しい公開テストではKimi K3のほうが優れていますが、公式の出力価格はGLM-5.2の3倍以上です。日常的なレビューを何千件も処理するチームなら、1ドルあたりの処理量ではGLMのほうが多くなる可能性があります。難しいビジュアルプロジェクトを1件救いたい開発者なら、K3のために喜んで料金を支払うでしょう。

Tiffany Layne | 2026-07-28

Kimi K3 対 GPT-5.6 Sol:トークンが安いのか、それともタスクが安いのか?

Kimi K3 対 GPT-5.6 Sol:トークンが安いのか、それともタスクが安いのか?

TL;DR 更新 — 2026年7月28日 :Kimi K3の完全な重みが公開されました。Moonshot AIは、公式リポジトリで2.8Tチェックポイント、技術レポート、Kimi K3 Licenseを公開しました。このリリースにより、GPT-5.6 Solに対するK3の制御性とデプロイ面での優位性は高まりましたが、独立ベンチマークの結果が変わったわけではなく、K3を自前で運用するコストが下がったわけでもありません。 Kimi K3はトークン単価が安く、GPT-5.6 Solは重要な本番エージェントのデフォルトとして優れています。この2つは両立します。 価格表が示すほど差は大きくありません。Artificial Analysisのテストでは、GPT-5.6 Sol maxのIntelligence Indexは59で、Kimi K3は57です。一方、測定されたタスクあたりのコストはSolが約1.04ドル、K3が0.95ドルで、公式の出力価格から想像される2倍の差ではありません。 簡単に言えば、幅広い信頼性、コーディングエージェントの性能、OpenAIのホスト型ツール群を重視するなら GPT-5.6 Sol を選びます。動画入力、長文コンテキスト、低い定価、または公開されたオープンウェイトへのアクセスが判断を左右するなら Kimi K3 を選びます。

Schuyler Stacy | 2026-07-28

Qwen 3.8 Maxとは?リリース日、2.4Tプレビュー、料金、初期ベンチマーク

Qwen 3.8 Maxとは?リリース日、2.4Tプレビュー、料金、初期ベンチマーク

Qwen 3.8 Maxは、Alibabaが開発した2.4兆パラメーターの新しい中国製フラッグシップモデルです。ただし、現在利用できるバージョンはまだ qwen3.8-max-preview と呼ばれています。このサフィックスは重要です。2026年7月21日時点で、Alibabaは本番モデル、テクニカルレポート、通常のトークン単価によるAPI料金、またはダウンロード可能な重みを公開していません。 プレビューは7月19日、AlibabaのToken Plan、Qoder、QoderWorkを通じて公開されました。 公式Qwen発表 で、チームはQwen 3.8が「近日中に」オープンウェイトになると述べ、Fable 5に次ぐ、最先端の主要モデルに匹敵する性能を持つと説明しています。ただし、その順位を裏付ける公開ベンチマークスイートや評価方法は示されていません。 私の簡単な見解:Qwen 3.8 Maxは実在する非常に興味深いプレビューですが、完成した製品の正式リリースではありません。開発者はテストを行い、すべての結果に日付を記録し、Alibabaがまだ公開していない仕様を前提に本番環境への移行計画を立てることは避けるべきです。

Tiffany Layne | 2026-07-23