什麼是 OpenAI 最新的 Astra 模型?發布日期、基準測試與比較(2026)

什麼是 OpenAI 最新的 Astra 模型?了解其發布日期、定價與基準測試的已確認事實,以及它與 GPT-5.6、Kimi K3 和 Fable 5 的比較。

什麼是 OpenAI 最新的 Astra 模型?發布日期、基準測試與比較(2026)

快速回答: OpenAI 最新的模型暫定名為 Astra,是一個研究階段的 多代理 AI 系統,於 2026 年 8 月 1 日首次亮相。Astra 不會一次完成回答,而是將問題拆分成多個部分,並在數小時或數天內協調一組子代理。其最受矚目的成果是:內部版本 解決了十個懸而未決的數學與理論計算機科學問題,這些問題至少十年來一直無法由人類研究人員解決,而總計算成本約為 $2,000。截至本文撰寫時,Astra 尚無公開發布日期,也沒有公布定價.

本指南將回答人們實際搜尋的問題:

  • Astra 是什麼?它究竟做了什麼?

  • Astra 何時發布?現在可以使用嗎?

  • Astra 的定價是多少?

  • Astra 是 OpenAI 的新旗艦模型,還是 GPT-6?

  • Astra 與 GPT-5.6、Kimi K3 和 Fable 5 相比如何?

  • 有哪些風險與未解答的問題?

全文會標示 ✅ 已確認 的事實,以及 ⚠️ 傳聞/未驗證 的說法,幫助你分辨可靠資訊與炒作。

目錄

What Is OpenAI Astra?

✅ Confirmed: Astra is a research-stage multi-agent system from OpenAI. Rather than producing a one-shot reply like a traditional chatbot, a root agent decomposes a task into sub-problems, spins up sub-agents to work on each piece, waits for their results, and synthesizes a final answer.

That architecture — not a benchmark score — is the real announcement. Frontier models until now have competed on how well a single model answers in one sitting. Astra is best understood as a coordination layer rather than a bigger brain. Its defining capability is managing a team of AI agents across long-horizon work — a single objective can run for hours or even days.

The multi-agent architecture, in plain terms

Think of Astra less as a smarter individual and more as a project manager with a research team. The root agent plans; the sub-agents execute in parallel; the root agent reviews and combines. This is why OpenAI frames Astra as built for persistent, extended tasks rather than quick Q&A.

Why it's not "just a bigger chatbot"

Most model launches lead with a demo video or a benchmark table. Astra's launch led with mathematics — actual research results. That framing signals OpenAI's bet: the next frontier isn't answering one question faster, it's holding a coherent line of reasoning across a long, multi-step task without drifting.


What Did Astra Actually Do? The 10 Math Problems

✅ Confirmed: An internal version of Astra produced solutions to ten open problems spanning group theory, coding theory, quantum complexity, high-dimensional geometry, lattice cryptography, arithmetic circuit complexity, and extremal combinatorics. All ten had been open for at least a decade — most far longer.

Non-sofic groups and Gromov's 1999 question

The standout result is in group theory: a construction proving that non-sofic groups exist. Whether every group is sofic (approximable by finite permutations) was a question posed by mathematician Mikhail Gromov in 1999 and had stayed open ever since. Astra also contributed to a disproof of Connes's rigidity conjecture. These are genuine research contributions, not exam scores.

Lean-verified proofs — why this matters

Two details separate this from earlier "AI does math" headlines:

  1. The proofs are machine-checkable. Every argument was formalized in Lean 4, a proof assistant that verifies each logical step. A Lean-certified proof doesn't depend on trusting the model — it can be independently re-run and checked.

  2. Humans wrote the papers. OpenAI has been explicit: the model generated the mathematical arguments, while humans handled authorship and exposition.

Outside researchers took it seriously. Thomas Bloom of the University of Manchester called the batch "big news." And the economics are striking — all ten solutions reportedly cost about $2,000 at API rates, suggesting the limiting factor is no longer compute price but the ability to steer a model across a long task.

Astra Release Date & Preview Status

When is the Astra release date? ✅ Confirmed: There is no public release date. The system that solved the ten problems is an internal research version. OpenAI previewed Astra on August 1, 2026, followed with a paper and machine-checkable proofs on August 6, and Sam Altman demonstrated it to policymakers in Washington, D.C. — but the model remains in testing.

Notably, OpenAI has signaled that any public release would land under the new U.S. federal AI framework, meaning Astra could be the first model to require federal approval before launch. That regulatory gate could keep Astra "in the display case" for months even if it's technically ready.

⚠️ Rumor / Unverified: A single social-media source (@synthwavedd, Aug 6) claimed Astra is a brand-new pretrain — the largest since GPT-4.5 — that the latest internal checkpoint is codenamed "mewfour," and that a launch could come "as early as the following week." These timing and scale claims are unverified, and prediction markets assigned low odds to a mid-August public release. Treat any "Astra is out now" headline with skepticism.

Astra Pricing — What We Know (and Don't)

What is Astra's pricing? ✅ Confirmed: OpenAI has announced no pricing, token rates, or subscription tier for Astra. Any total-cost-of-ownership estimate is speculative until OpenAI publishes numbers.

The only cost data point is that reference ~$2,000 figure for solving all ten math problems, computed at GPT-5.6 Sol API rates. That's a compute-cost anecdote, not a price list — useful for intuition, not budgeting.

Bottom line for teams: you cannot deploy Astra today, and you cannot price it. If you need a capable frontier model right now, you'll use something that's already available (more on that below).

Is Astra OpenAI's New Flagship? Is It GPT-6?

✅ Confirmed, with caveats: OpenAI is positioning Astra as its next major model family — not an incremental update. But it has not decided whether Astra ships as GPT-6 or as a variant within the GPT-5 line, and "Astra" is itself a tentative codename for the model class.

So today, the publicly available flagship remains GPT-5.6 Sol. Astra is the signpost for where OpenAI is heading — an "upgrade" in ambition and architecture — but it is not yet a product you can select in a dropdown.

Astra Benchmarks & How It Compares

Are there Astra benchmarks? ✅ Confirmed: OpenAI deliberately did not release a benchmark table. Instead of a leaderboard, it showed original research results in fields where no benchmark even exists. That's a strategic move — it reframes competition from "higher exam score" to "can you produce new knowledge."

That makes direct head-to-head numbers impossible for now. What we can do is place Astra in context against the models you can actually use today. Here's how the landscape looks:

Astra vs GPT-5.6 vs Kimi K3 vs Fable 5 — comparison matrix

Dimension OpenAI Astra GPT-5.6 Sol Kimi K3 Claude Fable 5
Status ⚠️ Research preview, not released ✅ Available now ✅ Available (open weights) ✅ Available now
Type Multi-agent coordination system Single flagship LLM 2.8T-param open MoE model Closed frontier flagship
Best at Long-horizon research, original proofs General agentic + reasoning Coding, agents, cost-efficiency Top-tier reasoning & real-world tasks
Signature result 10 decade-old math problems, Lean-verified Established production flagship Beats most models on coding benchmarks Leads knowledge-work benchmarks
Pricing (per 1M tokens) ⚠️ Unknown ~$4 in / $24 out ~$2.7 in / $13.5 out ~$8 in / $40 out
You can use it today? ❌ No ✅ Yes ✅ Yes ✅ Yes

Pricing above reflects aggregated API rates (e.g., via GPT Proto); official direct rates may differ.

Astra vs GPT-5.6 Sol

GPT-5.6 Sol is OpenAI's current public flagship and already supports agentic workflows and sub-agents. Astra is the dedicated next step for persistent, multi-day work. If you need production reliability today, Sol is the answer; Astra is the future direction.

Astra vs Kimi K3

Kimi K3 (Moonshot AI) is a 2.8-trillion-parameter open model that punches far above its price — strong on coding and agent benchmarks at roughly a third of flagship costs. Where Astra targets frontier research, Kimi K3 targets practical, high-volume engineering work at low cost. Different jobs entirely, but for most builders K3 is usable now and cheap.

Astra vs Fable 5

Claude Fable 5 is a top-tier closed flagship, frequently leading real-world knowledge-work and reasoning evaluations. It's the premium option you can actually deploy. Astra's pitch isn't "beat Fable 5 on a benchmark" — it's "produce results in domains where benchmarks don't exist." Until Astra ships, Fable 5 is the higher-end available choice.

Astra Risks & Open Questions

Before you get swept up in the hype, weigh these ⚠️ risks and caveats:

  • Cherry-picked demo ≠ reliability. Ten hand-selected, Lean-verified proofs are a serious signal — but they say little about how Astra performs on a random problem chosen by an outside researcher. The real test comes when strangers can run it on their own questions.

  • Regulatory gate. As the first model potentially subject to a new U.S. federal approval framework, Astra's release could be delayed for months regardless of technical readiness.

  • Rumor contamination. Much of the "Astra" coverage — especially the "launching next week" and "mewfour" claims — traces to a single unverified source. Some outlets have reported speculation as fact.

  • No reproducibility yet. Whether Astra's approach generalizes beyond a curated set is the open question. A capability that only works on hand-picked problems is a demo; one that works on a stranger's question is a product.

Should You Wait for Astra — or Use a Model Today?

Here's the honest take: Astra is not something you can use. No release date, no pricing, no API. If your work is frontier mathematics research, Astra is worth watching. For everyone else — building apps, writing code, running agents, generating content — the smart move is to use the models that are already shipping.

The good news: the exact models Astra will be compared against — GPT-5.6 Sol, Kimi K3, and Claude Fable 5 — are all available right now, and you don't need three separate accounts to try them. On an aggregator like GPT Proto's model library, a single API key and one shared balance unlocks 200+ models across text, image, video, and audio — including GPT-5.6, Kimi K3, and Fable 5 — often at 10–60% below official rates.

That means you can benchmark Astra's future rivals against each other today, pick the best fit for your task, and swap models without rewiring your stack. When Astra eventually ships, you'll already have the infrastructure to add it in one line.

👉 Compare and try GPT-5.6, Kimi K3, Fable 5 and 200+ models on GPT Proto →

常見問題

OpenAI 最新的模型 Astra 是什麼?

Astra 是 OpenAI 的研究階段多代理 AI 系統,於 2026 年 8 月 1 日首次亮相。它能協調子代理處理長時間跨度的任務,並以約 2,000 美元的計算成本,透過 Lean 驗證的證明解決十個有數十年歷史的數學問題。

Astra 何時發布?

目前沒有公布發布日期。Astra 仍是處於測試中的內部研究版本,可能成為首個在公開發布前需要美國聯邦批准的模型。關於「下週」即將發布的報導尚未獲得驗證。

Astra 的費用是多少?

OpenAI 尚未公布任何定價。唯一公開的成本數字是解決全部十個數學問題約需 2,000 美元,按 GPT-5.6 Sol API 費率計算;這是計算成本案例,而非價格表。

Astra 是 GPT-6 嗎?

尚未確認。OpenAI 尚未決定 Astra 將以 GPT-6 形式發布,還是作為 GPT-5 系列的變體。「Astra」是下一個主要模型系列的暫定代號。

Astra 真的解決了數學問題嗎?

是。Astra 的內部版本為十個未解決問題提出了論證,之後以 Lean 4 形式化,其中包括首次明確構造非 sofic 群(回答 Gromov 的 1999 年問題),以及否證 Connes 的剛性猜想。

Astra 與 GPT-5.6、Kimi K3 和 Fable 5 相比如何?

Astra 瞄準前沿研究與長時間跨度的多代理工作,但目前尚未公開提供。GPT-5.6 Sol、Kimi K3 與 Fable 5 現在都可以使用——Sol 是 OpenAI 目前的旗艦模型,Kimi K3 是低成本的開放式強大模型,而 Fable 5 是高階封閉式前沿模型。

我現在可以預覽或使用 Astra 嗎?

不行。目前沒有公開預覽、API 或候補名單。若要與可比較的前沿模型合作,可使用像 [GPTProto](https://gptproto.com/model), 的聚合平台,透過一組 API 金鑰提供 GPT-5.6、Kimi K3、Fable 5 及其他 200 多個模型。

Astra 是風險還是被過度炒作?

Lean 驗證的證明是真正的成就,但展示經過精心挑選,在隨機問題上的可靠性尚未證實,而且許多「發布」相關傳聞也未經驗證。請將 Astra 視為有前景的訊號,而不是已推出的產品。

相關文章

更多部落格
2026 年 7 款最實惠的程式設計 LLM:API 價格與效能比較

2026 年 7 款最實惠的程式設計 LLM:API 價格與效能比較

最便宜的程式設計模型,不一定是使用成本最低的模型。 每百萬個輸入 token 價格僅 0.14 美元的模型看似便宜,但如果它誤解程式碼儲存庫、修改錯誤檔案,還需要重試三次,實際成本就不一定最低。另一方面,token 價格較高的模型,可能一次就能完成相同的修補。 因此,這不是另一份單純依輸入價格排序的模型清單。 我們首先尋找具備足夠程式設計能力的模型,確保它們能處理終端機操作、除錯與多步驟開發任務。接著,我們使用兩種相同的模擬工作負載,比較它們的輸入、快取輸入與輸出價格。 本排名涵蓋可透過 API 存取的 LLM ,不包含程式設計 IDE 訂閱服務。我們也排除了自行託管的模型,因為 GPU、推論基礎架構、維護與工程時間都不是免費的。 價格與基準測試結果已於 2026 年 8 月 12 日 核對。請將這些資料視為當時的快照,而不是永久適用的價目表。

Michael Johnson | 2026-08-12

2026 年 7 款最佳中國 AI 影片模型,依實際使用情境比較

2026 年 7 款最佳中國 AI 影片模型,依實際使用情境比較

Last checked: August 2026 Chinese AI video models are no longer simply cheaper alternatives to products from the United States. Seedance, MiniMax, Wan, Kling, Vidu, and other Chinese video AI model families now compete at the top of independent leaderboards, while introducing features such as 30-second generation, native audio, video editing, multiple reference assets, and even document-to-video creation. The difficult part is knowing what you are actually comparing. Dreamina is not a model, Hailuo and MiniMax H3 are not interchangeable names, and Qwen is not Alibaba's primary video-generation family. A model with the highest advertised resolution may also be the wrong choice for character acting, product consistency, or high-volume image-to-video work. This guide compares seven of the best Chinese AI video models in 2026 by what each one is genuinely best suited to do. It also separates independently verified performance from newly announced capabilities that still need broader testing. Want to try several models before committing to one? Explore the AI video generation workspace to compare available text-to-video, image-to-video, and reference-to-video models in one place.

Schuyler Stacy | 2026-08-11

GLM 5.2 與 Claude Opus 5:哪個程式編碼模型更具成本效益?

GLM 5.2 與 Claude Opus 5:哪個程式編碼模型更具成本效益?

低廉的 token 不一定能帶來低成本的結果。在 GLM 5.2 與 Opus 5 的比較中,這項區別格外重要,因為表面數據指向相反方向:GLM-5.2 成本較低、回應速度較快,而 Claude Opus 5 在目前獨立智慧比較中領先,且除了文字之外也能檢視圖片。 我的簡短答案很直接。對於大量、範圍明確的編碼工作,且結果會由開發人員或更強的審查模型檢查時,選擇 GLM-5.2。對於模糊的儲存庫變更、視覺化前端偵錯,以及首次嘗試失敗的成本高於模型呼叫成本的任務,選擇 Claude Opus 5。 有一個原因讓我們必須謹慎看待更強的結論。Z.ai 於 2026 年 6 月發布 GLM-5.2,但 Anthropic 於 7 月 24 日發布 Opus 5。大多數社群討論與「真實世界」比較仍然是以 GLM-5.2 對比 Opus 4.8。這些結果可作為有用的背景資料,但不能證明 GLM-5.2 勝過或不如 Opus 5。 本文是以證據為基礎的比較,而非第一手基準測試。結論來自目前的模型文件、GPTProto 定價、獨立基準資料、供應商披露資訊,以及社群評估方法。若目前尚無直接的 GLM-5.2 與 Opus 5 證據,我們會明確說明這項限制。

Michael Johnson | 2026-08-04

GLM-5.2 與 Kimi K3 程式設計比較:2026 年哪個更適合開發者?

GLM-5.2 與 Kimi K3 程式設計比較:2026 年哪個更適合開發者?

TL;DR: 當任務困難、執行時間長或涉及視覺內容時,Kimi K3 是更強的程式設計模型。在 Moonshot 公開的程式設計比較中,它全面領先 GLM-5.2,並可透過其託管服務接受圖片與影片。對於日常的儲存庫工作,GLM-5.2 仍是更好的預設選擇:成本低得多、運行規模較小,且採用寬鬆的 MIT 授權。Kimi K3 現在也已釋出權重,但其 1.56 TB 儲存庫、建議使用 64 個以上加速器的部署要求,以及自訂授權,意味著自行託管需要投入更多資源。當能力是瓶頸時選擇 Kimi;當成本與日常運營簡易性更重要時選擇 GLM。 GLM-5.2 與 Kimi K3 程式碼比較中有趣的地方,不在於兩個模型都能撰寫 React 元件或解決簡短演算法。這個層級的模型早已具備這些能力。真正有用的問題是,當任務變得複雜時會發生什麼:儲存庫稽核、多檔案遷移、只會在螢幕截圖中出現的錯誤,或必須讓多個系統保持一致的可遊玩 Three.js 原型。 這也是價格差異開始產生影響的地方。Kimi K3 在最困難的公開測試中表現較佳,但其官方輸出價格超過 GLM-5.2 的三倍。每天執行數千次普通審查的團隊,使用 GLM 可能能以每美元完成更多工作。試圖挽救一個棘手視覺專案的開發者,則可能很樂意為 K3 買單。

Tiffany Layne | 2026-07-28