兩款前沿模型在四週內相繼發布,目標買家也高度重疊:執行代理程式、而非聊天機器人的開發者。Moonshot AI 於 2026 年 7 月 16 日推出 Kimi K3,xAI 則於 8 月 12 日以 Grok 4.6 回應。今天搜尋 "Grok 4.6 vs Kimi K3",您會看到雙方陣營的發布報導,以及大量規格表——但幾乎沒有人從建構者的角度將兩者並列比較。這正是本文要填補的空白。
以下是簡短版,因為您是來做決策,而不是看回顧。
Grok 4.6 在代理式執行效率與免操心託管方面勝出。 它能以更少的迴圈和更少的 token 完成長時間、多步驟的任務,而且您完全不必接觸基礎架構。 Kimi K3 則在上下文、原生影片支援與控制權方面勝出——具備 100 萬 token 的上下文視窗、影像 and 影片輸入,以及可下載的開放權重;如果您需要自行託管或隔離網路環境,這些特點非常重要。在大家都引用的那個數字上,兩者幾乎打成平手:Artificial Analysis 將兩者的單任務成本都估在約 $0.84。因此, intelligence index 只差一分並不是您做決定的關鍵。兩個模型以相反的路徑達到相同的成本,而 that 正是您真正需要選擇的分岔點。
如果您執行的是重視成本、高流量的代理工作流程,並希望使用託管端點,請選 Grok 4.6。如果您需要將整個儲存庫或影片放入單一上下文視窗,或因合規要求而必須自行持有權重,請選 Kimi K3。本文其餘內容將展示這項判斷背後的依據。
Specs and pricing, side by side
A few notes before the table. Benchmark figures published at launch come from each vendor's own materials; treat them as 🟡 vendor-reported until independent runs confirm them. The Artificial Analysis numbers (intelligence index, per-task cost) are the closest thing to a neutral head-to-head and are marked as such. The GPT Proto rate is what you actually pay to test both through one key; the official list is the vendor's own price.
|
Grok 4.6 |
Kimi K3 |
| Vendor |
xAI |
Moonshot AI |
| Released |
Aug 12, 2026 |
Jul 16, 2026 |
| Architecture |
Grok 4.5 family, extended supplemental training + agentic RL |
2.8T-param MoE, 16 of 896 experts active per token |
| Model access |
Hosted API only (no downloadable weights) |
Open-weight (weights released Jul 27, 2026, under a custom Kimi K3 License) |
| Context window |
500,000 tokens |
1,048,576 tokens |
| Inputs |
Text, image |
Text, image, video |
| Output |
Text (no fixed output cap) |
Text; up to 1,048,576 within the context budget, 131,072 default |
| AA Intelligence Index 🟡 |
61 |
~57 (Artificial Analysis); vendor charts place it just under Grok |
| Per-task cost (Artificial Analysis) |
~$0.84 |
~$0.84 |
| Official list price /1M tokens |
$2 in / $6 out (short context) |
$3 in / $15 out |
| GPT Proto price /1M tokens |
3.60 out (40% off) |
13.50 out (10% off) |
The pricing rows deserve a second look, because the headline "Grok is cheaper" is true but incomplete. Grok 4.6's $2/$6 list rate only holds under 200K tokens. Cross that line — which a large-codebase agent does routinely — and the rate doubles to $4/$12, applied to the entire request, not just the overflow. Kimi K3 charges more per output token, but its full 1M context carries no surcharge. So the cheaper model flips depending on how long your prompts run. Short, high-frequency calls favor Grok; long-document or repository-scale work narrows the gap and can invert it.
Intelligence and knowledge work
On the composite Artificial Analysis Intelligence Index — nine benchmarks compressed into one digit — Grok 4.6 scores 61 and Kimi K3 sits a few points below. I'd flag two things about that gap. First, the exact Kimi number wanders by source: Artificial Analysis reports 57, some launch coverage cites 60, and xAI's own chart just says "higher than Kimi." I'm anchoring to the Artificial Analysis figure because it's the one neutral run, but a one-to-four-point spread on a nine-benchmark composite is well inside the noise of how you weight the components. Second — and this matters more — a composite score measures the quality of an answer. It barely measures whether a model can hold a goal across forty tool calls without losing the plot. That failure mode is the one that actually kills agent deployments, and neither vendor's index number predicts it.
Where they genuinely diverge is agentic efficiency. On Artificial Analysis's private AA-Briefcase workload, Grok 4.6 finishes in about 53 turns and roughly 0.5 billion input tokens. For scale: Claude Opus 5 needs around 103 turns and 2.0 billion tokens on the same task. That's the concrete payoff of xAI's "self-testing on long trajectories" — the model checks its own work before taking the next step, so it loops less. Kimi K3 is competitive on the outcome of long knowledge work (it lands in the Fable 5 tier on AA-Briefcase Elo, ~1548) but it doesn't advertise the same turn-count discipline. If your bill scales with tool calls — and in production agents, it does — this is a real, measurable difference, not a marketing line.
The takeaway in plain words: they're roughly matched on how smart the answers are, but Grok 4.6 gets there in fewer moves.
Coding: it depends which coding
This is where a blanket winner falls apart, so here are the specifics. Kimi K3 leads on several coding suites: SWE Marathon (42.0 vs. reference frontier scores of 35–40), Program Bench (77.8, edging GPT-5.6 Sol's 77.6), and it lands within half a point of the top on Terminal-Bench 2.1 at 88.3 (tracked in Artificial Analysis's Coding Agent Index). On the web-browsing benchmark BrowseComp it tops the field at 91.2. Grok 4.6, meanwhile, posted 88.4% on Terminal-Bench v2.1 and a 1753 Elo on GDPval-AA v2 — and its sharpest gains over Grok 4.5 show up on CursorBench 3.2 (69.9% vs. 66.7%) and Terminal-Bench v3.0 (26% vs. 15.7%).
Read past the numbers and a pattern emerges. Kimi K3's coding strength is broad and repository-shaped: it navigates large codebases, debugs from logs and screenshots, and its 1M window lets it hold an entire project in context. Grok 4.6's coding strength is trajectory-shaped: it was tuned inside Cursor and Grok Build against real developer sessions, so it excels at sustaining a coding task — turning a vague idea into a running first version and refining it across many steps. The cost, and every model has one here: Grok's 500K window caps how much of a monorepo it can see at once, and it can't hold video feedback. Kimi's edge on raw coding benchmarks comes with a higher output-token bill and, per Moonshot's own launch note, a hallucination rate that ticked up alongside the accuracy gains.
For frontend and interactive work specifically — a common variant of this search — Kimi K3 topped LMArena's Frontend Code Arena at launch, while Grok 4.6 landed near GPT-5.6 Sol and Claude Fable on Code Arena web-dev tasks. Both are strong; Kimi has the sharper independent signal on frontend today.
Context, multimodal, and deployment
Here the two models aren't competing on a spectrum — they're built differently. Kimi K3 gives you a 1,048,576-token context window and native input for text, images, and video, from a single architecture. That combination is the reason to reach for it: a coding agent that reads screenshots to refine a UI, a research workflow that keeps a full document corpus resident, a QA system that compares interface video against implementation. Grok 4.6 offers 500K tokens and text-plus-image input. Ample for most work, but if your job is "analyze this 40-minute screen recording" or "hold these 300 files in one prompt," Grok isn't the tool.
Then there's the ownership question, which is binary. Grok 4.6 is hosted-only; there is nothing to download, and xAI can revise or deprecate the endpoint underneath you. Kimi K3 ships open weights under a custom Kimi K3 License — you can self-host, fine-tune, and quantize. For teams with data-residency rules, air-gap requirements, or a policy against vendor lock-in, that's decisive. The honest caveat: at native 4-bit precision the weights need roughly 1.4 TB resident before any KV cache, which exceeds a single 8-GPU node. "Open" here means legally and technically available to those with serious hardware, not "runs on your laptop." For everyone else, the practical path to Kimi K3 is a hosted API anyway.
When to pick which
No fence-sitting. Here's the call by workload.
Pick Kimi K3 if any of these describe you: you need to feed an entire repository, a long document set, or video into one context window; you're doing frontend or interactive coding where its LMArena lead shows; you have a compliance or lock-in reason to hold the weights; or you want native multimodal without stitching a separate vision model on. You'll pay more per output token — budget for that on high-volume text generation.
Pick Grok 4.6 if: you run long-horizon autonomous agents and care about finishing in fewer turns and fewer tokens; your prompts stay under 200K so you get the clean $2/$6 (list) rate; you want a managed endpoint with zero infrastructure; or you're cost-sensitive at high frequency. Watch the long-context tier — cross 200K and the whole request bills at double.
The tie on per-task cost is the point. You're not choosing a cheaper model. You're choosing between hosted turn-efficiency and open long-context multimodality. Match that to your workload, not to the one-point index gap.
Run both on one key with GPT Proto
The fastest way to settle a "which is better for my task" argument is to run your own task on both. GPT Proto exposes Kimi K3 and Grok 4.6 through one OpenAI-compatible endpoint and one balance, so you can A/B them without funding two accounts. Note the auth header: GPT Proto's /v1/ surface takes the raw key with no Bearer prefix.
Python:
import openai
client = openai.OpenAI(
api_key="YOUR_GPTPROTO_API_KEY",
base_url="https://gptproto.com/v1",
)
prompt = "Refactor this function for readability and explain each change:\n\n<paste code>"
for model in ["kimi-k3", "grok-4.6"]:
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
)
print(f"=== {model} ===")
print(resp.choices[0].message.content)
print("tokens:", resp.usage.total_tokens)
cURL, first call:
curl https://gptproto.com/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: YOUR_GPTPROTO_API_KEY" \
-d '{
"model": "kimi-k3",
"messages": [
{"role": "user", "content": "Summarize the tradeoffs between MoE and dense LLMs in 3 bullets."}
]
}'
Two things specific to Kimi K3 that a plain model-name swap will miss. It always reasons — set reasoning_effort to low, high, or max (default max) at the top level, not the older thinking config. And in multi-turn or tool loops, return the complete previous assistant message, including reasoning_content, or you'll break reasoning continuity on long sessions. Parse your final answer from content, never from reasoning_content.
Full model details and current pricing: Kimi K3 on GPT Proto and Grok 4.6 on GPT Proto.