OpenAIの最新モデルAstraとは? リリース日・ベンチマーク・他モデルとの比較(2026)

OpenAIの最新モデルAstraとは? リリース日、料金、ベンチマークに関する確認済みの事実に加え、GPT-5.6、Kimi K3、Fable 5との比較も紹介します。

OpenAIの最新モデルAstraとは? リリース日・ベンチマーク・他モデルとの比較(2026)

Quick answer: OpenAI's newest model, tentatively named Astra, is a research-stage multi-agent AI system previewed on August 1, 2026. Instead of answering in a single pass, Astra breaks a problem into pieces and coordinates a team of sub-agents over hours or days. Its headline achievement: an internal version solved ten open math and theoretical–computer-science problems that had resisted human researchers for at least a decade — at a total compute cost of roughly $2,000. As of this writing, Astra has no public release date and no announced pricing.

This guide answers the questions people actually search for:

  • What is Astra, and what did it really do?

  • When is the Astra release date, and can you use it now?

  • What is Astra's pricing?

  • Is Astra OpenAI's new flagship — or GPT-6?

  • How does Astra compare to GPT-5.6, Kimi K3, and Fable 5?

  • What are the risks and open questions?

Throughout, we mark ✅ Confirmed facts and ⚠️ Rumor / Unverified claims so you can tell the signal from the hype.

目次

What Is OpenAI Astra?

✅ Confirmed: Astra is a research-stage multi-agent system from OpenAI. Rather than producing a one-shot reply like a traditional chatbot, a root agent decomposes a task into sub-problems, spins up sub-agents to work on each piece, waits for their results, and synthesizes a final answer.

That architecture — not a benchmark score — is the real announcement. Frontier models until now have competed on how well a single model answers in one sitting. Astra is best understood as a coordination layer rather than a bigger brain. Its defining capability is managing a team of AI agents across long-horizon work — a single objective can run for hours or even days.

The multi-agent architecture, in plain terms

Think of Astra less as a smarter individual and more as a project manager with a research team. The root agent plans; the sub-agents execute in parallel; the root agent reviews and combines. This is why OpenAI frames Astra as built for persistent, extended tasks rather than quick Q&A.

Why it's not "just a bigger chatbot"

Most model launches lead with a demo video or a benchmark table. Astra's launch led with mathematics — actual research results. That framing signals OpenAI's bet: the next frontier isn't answering one question faster, it's holding a coherent line of reasoning across a long, multi-step task without drifting.


What Did Astra Actually Do? The 10 Math Problems

✅ Confirmed: An internal version of Astra produced solutions to ten open problems spanning group theory, coding theory, quantum complexity, high-dimensional geometry, lattice cryptography, arithmetic circuit complexity, and extremal combinatorics. All ten had been open for at least a decade — most far longer.

Non-sofic groups and Gromov's 1999 question

The standout result is in group theory: a construction proving that non-sofic groups exist. Whether every group is sofic (approximable by finite permutations) was a question posed by mathematician Mikhail Gromov in 1999 and had stayed open ever since. Astra also contributed to a disproof of Connes's rigidity conjecture. These are genuine research contributions, not exam scores.

Lean-verified proofs — why this matters

Two details separate this from earlier "AI does math" headlines:

  1. The proofs are machine-checkable. Every argument was formalized in Lean 4, a proof assistant that verifies each logical step. A Lean-certified proof doesn't depend on trusting the model — it can be independently re-run and checked.

  2. Humans wrote the papers. OpenAI has been explicit: the model generated the mathematical arguments, while humans handled authorship and exposition.

Outside researchers took it seriously. Thomas Bloom of the University of Manchester called the batch "big news." And the economics are striking — all ten solutions reportedly cost about $2,000 at API rates, suggesting the limiting factor is no longer compute price but the ability to steer a model across a long task.

Astra Release Date & Preview Status

When is the Astra release date? ✅ Confirmed: There is no public release date. The system that solved the ten problems is an internal research version. OpenAI previewed Astra on August 1, 2026, followed with a paper and machine-checkable proofs on August 6, and Sam Altman demonstrated it to policymakers in Washington, D.C. — but the model remains in testing.

Notably, OpenAI has signaled that any public release would land under the new U.S. federal AI framework, meaning Astra could be the first model to require federal approval before launch. That regulatory gate could keep Astra "in the display case" for months even if it's technically ready.

⚠️ Rumor / Unverified: A single social-media source (@synthwavedd, Aug 6) claimed Astra is a brand-new pretrain — the largest since GPT-4.5 — that the latest internal checkpoint is codenamed "mewfour," and that a launch could come "as early as the following week." These timing and scale claims are unverified, and prediction markets assigned low odds to a mid-August public release. Treat any "Astra is out now" headline with skepticism.

Astra Pricing — What We Know (and Don't)

What is Astra's pricing? ✅ Confirmed: OpenAI has announced no pricing, token rates, or subscription tier for Astra. Any total-cost-of-ownership estimate is speculative until OpenAI publishes numbers.

The only cost data point is that reference ~$2,000 figure for solving all ten math problems, computed at GPT-5.6 Sol API rates. That's a compute-cost anecdote, not a price list — useful for intuition, not budgeting.

Bottom line for teams: you cannot deploy Astra today, and you cannot price it. If you need a capable frontier model right now, you'll use something that's already available (more on that below).

Is Astra OpenAI's New Flagship? Is It GPT-6?

✅ Confirmed, with caveats: OpenAI is positioning Astra as its next major model family — not an incremental update. But it has not decided whether Astra ships as GPT-6 or as a variant within the GPT-5 line, and "Astra" is itself a tentative codename for the model class.

So today, the publicly available flagship remains GPT-5.6 Sol. Astra is the signpost for where OpenAI is heading — an "upgrade" in ambition and architecture — but it is not yet a product you can select in a dropdown.

Astra Benchmarks & How It Compares

Are there Astra benchmarks? ✅ Confirmed: OpenAI deliberately did not release a benchmark table. Instead of a leaderboard, it showed original research results in fields where no benchmark even exists. That's a strategic move — it reframes competition from "higher exam score" to "can you produce new knowledge."

That makes direct head-to-head numbers impossible for now. What we can do is place Astra in context against the models you can actually use today. Here's how the landscape looks:

Astra vs GPT-5.6 vs Kimi K3 vs Fable 5 — comparison matrix

Dimension OpenAI Astra GPT-5.6 Sol Kimi K3 Claude Fable 5
Status ⚠️ Research preview, not released ✅ Available now ✅ Available (open weights) ✅ Available now
Type Multi-agent coordination system Single flagship LLM 2.8T-param open MoE model Closed frontier flagship
Best at Long-horizon research, original proofs General agentic + reasoning Coding, agents, cost-efficiency Top-tier reasoning & real-world tasks
Signature result 10 decade-old math problems, Lean-verified Established production flagship Beats most models on coding benchmarks Leads knowledge-work benchmarks
Pricing (per 1M tokens) ⚠️ Unknown ~$4 in / $24 out ~$2.7 in / $13.5 out ~$8 in / $40 out
You can use it today? ❌ No ✅ Yes ✅ Yes ✅ Yes

Pricing above reflects aggregated API rates (e.g., via GPT Proto); official direct rates may differ.

Astra vs GPT-5.6 Sol

GPT-5.6 Sol is OpenAI's current public flagship and already supports agentic workflows and sub-agents. Astra is the dedicated next step for persistent, multi-day work. If you need production reliability today, Sol is the answer; Astra is the future direction.

Astra vs Kimi K3

Kimi K3 (Moonshot AI) is a 2.8-trillion-parameter open model that punches far above its price — strong on coding and agent benchmarks at roughly a third of flagship costs. Where Astra targets frontier research, Kimi K3 targets practical, high-volume engineering work at low cost. Different jobs entirely, but for most builders K3 is usable now and cheap.

Astra vs Fable 5

Claude Fable 5 is a top-tier closed flagship, frequently leading real-world knowledge-work and reasoning evaluations. It's the premium option you can actually deploy. Astra's pitch isn't "beat Fable 5 on a benchmark" — it's "produce results in domains where benchmarks don't exist." Until Astra ships, Fable 5 is the higher-end available choice.

Astra Risks & Open Questions

Before you get swept up in the hype, weigh these ⚠️ risks and caveats:

  • Cherry-picked demo ≠ reliability. Ten hand-selected, Lean-verified proofs are a serious signal — but they say little about how Astra performs on a random problem chosen by an outside researcher. The real test comes when strangers can run it on their own questions.

  • Regulatory gate. As the first model potentially subject to a new U.S. federal approval framework, Astra's release could be delayed for months regardless of technical readiness.

  • Rumor contamination. Much of the "Astra" coverage — especially the "launching next week" and "mewfour" claims — traces to a single unverified source. Some outlets have reported speculation as fact.

  • No reproducibility yet. Whether Astra's approach generalizes beyond a curated set is the open question. A capability that only works on hand-picked problems is a demo; one that works on a stranger's question is a product.

Should You Wait for Astra — or Use a Model Today?

Here's the honest take: Astra is not something you can use. No release date, no pricing, no API. If your work is frontier mathematics research, Astra is worth watching. For everyone else — building apps, writing code, running agents, generating content — the smart move is to use the models that are already shipping.

The good news: the exact models Astra will be compared against — GPT-5.6 Sol, Kimi K3, and Claude Fable 5 — are all available right now, and you don't need three separate accounts to try them. On an aggregator like GPT Proto's model library, a single API key and one shared balance unlocks 200+ models across text, image, video, and audio — including GPT-5.6, Kimi K3, and Fable 5 — often at 10–60% below official rates.

That means you can benchmark Astra's future rivals against each other today, pick the best fit for your task, and swap models without rewiring your stack. When Astra eventually ships, you'll already have the infrastructure to add it in one line.

👉 Compare and try GPT-5.6, Kimi K3, Fable 5 and 200+ models on GPT Proto →

よくある質問

OpenAIの最新モデルAstraとは何ですか?

AstraはOpenAIの研究段階にあるマルチエージェントAIシステムで、2026年8月1日にプレビュー公開されました。長期的なタスクにわたってサブエージェントを調整し、約2,000ドルの計算コストでLean検証済みの証明を用いて10件の10年来の数学問題を解いたことで知られています。

Astraはいつリリースされますか?

発表されたリリース日はありません。Astraはテスト中の内部研究用バージョンのままであり、一般公開前に米国連邦政府の承認を必要とする最初のモデルになる可能性があります。「来週リリース」という近い時期の報道は未検証です。

Astraの料金はいくらですか?

OpenAIは価格を一切公表していません。公表されている唯一のコスト数値は、10件すべての数学問題を解くのに約2,000ドル(GPT-5.6 SolのAPIレートで計算)というもので、これは価格表ではなく計算コストの事例です。

AstraはGPT-6ですか?

確認されていません。OpenAIはAstraをGPT-6として出すか、GPT-5系の変種として出すかを決定していません。「Astra」は次期メジャーモデルファミリーの暫定的なコードネームです。

Astraは本当に数学問題を解いたのですか?

はい。Astraの内部バージョンは、10の未解決問題に対する議論を生成し、後にLean 4で形式化されました。これには、非ソフィック群の最初の明示的構成(グロモフの1999年の問題への回答)と、コンヌの剛性予想の反証が含まれます。

AstraはGPT-5.6、Kimi K3、Fable 5とどう比較されますか?

Astraはフロンティアの*研究*と長期的なマルチエージェント作業を対象としていますが、一般公開されていません。GPT-5.6 Sol、Kimi K3、Fable 5はすべて今日利用可能です。SolはOpenAIの現在のフラッグシップ、Kimi K3は低コストなオープンの強豪、Fable 5はプレミアムなクローズドフロンティアモデルです。

今すぐAstraをプレビューまたは利用できますか?

いいえ。一般プレビュー、API、ウェイトリストはありません。現在、同等のフロンティアモデルを利用するには、[GPTProto](https://gptproto.com/model), のようなアグリゲーターを使いましょう。1つのAPIキーでGPT-5.6、Kimi K3、Fable 5、その他200以上のモデルを利用できます。

Astraはリスクですか、過大評価されていますか?

Leanで検証された証明は本物の成果ですが、デモは厳選されたものであり、ランダムな問題での信頼性は証明されておらず、「リリース」に関する情報の多くは未検証です。Astraは、出荷された製品ではなく、有望なシグナルとして扱ってください。
2026年版コーディング向け低価格LLMベスト7:API料金と性能の比較

2026年版コーディング向け低価格LLMベスト7:API料金と性能の比較

最も安価なコーディングモデルが、必ずしも最も安く使えるモデルとは限りません。 $0.14 / 100万入力トークンのモデルは安価に見えます。しかし、リポジトリを誤解し、間違ったファイルを編集し、3回の再試行が必要になるまではそう思えるでしょう。一方、トークン単価が高いモデルでも、同じパッチを1回で完了できる場合があります。 だからこそ、これは入力料金だけでモデルを並べた、よくあるランキングではありません。 まず、ターミナル操作、デバッグ、複数ステップの開発タスクに対応できる十分なコーディング能力を持つモデルを探しました。次に、同じ2種類のシミュレーションワークロードを使って、入力、キャッシュ入力、出力の料金を比較しました。 このランキングでは、コーディングIDEのサブスクリプションではなく、 APIから利用できるLLM を対象にしています。また、GPU、推論インフラ、保守、エンジニアリング時間は無料ではないため、セルフホスト型モデルは除外しています。 料金とベンチマーク結果は 2026年8月12日 時点で確認しました。恒久的な料金表ではなく、その時点のスナップショットとして扱ってください。

Michael Johnson | 2026-08-12

2026年中国のAI動画モデルベスト7、実際の用途別に比較

2026年中国のAI動画モデルベスト7、実際の用途別に比較

Last checked: August 2026 Chinese AI video models are no longer simply cheaper alternatives to products from the United States. Seedance, MiniMax, Wan, Kling, Vidu, and other Chinese video AI model families now compete at the top of independent leaderboards, while introducing features such as 30-second generation, native audio, video editing, multiple reference assets, and even document-to-video creation. The difficult part is knowing what you are actually comparing. Dreamina is not a model, Hailuo and MiniMax H3 are not interchangeable names, and Qwen is not Alibaba's primary video-generation family. A model with the highest advertised resolution may also be the wrong choice for character acting, product consistency, or high-volume image-to-video work. This guide compares seven of the best Chinese AI video models in 2026 by what each one is genuinely best suited to do. It also separates independently verified performance from newly announced capabilities that still need broader testing. Want to try several models before committing to one? Explore the AI video generation workspace to compare available text-to-video, image-to-video, and reference-to-video models in one place.

Schuyler Stacy | 2026-08-11

GLM 5.2 vs Claude Opus 5:どちらのコーディングモデルがより費用対効果に優れているか?

GLM 5.2 vs Claude Opus 5:どちらのコーディングモデルがより費用対効果に優れているか?

安価なトークンが、必ずしも安価な結果を意味するわけではありません。GLM 5.2とOpus 5の比較では、この違いが重要です。見出しの数字が正反対の方向を示しているからです。GLM-5.2は低価格で応答も高速ですが、Claude Opus 5は現在の独立した知能比較で首位に立ち、テキストだけでなく画像も検査できます。 結論を先に言うと、シンプルです。開発者またはより強力なレビュー用モデルが結果を確認する、高ボリュームで範囲の明確なコーディング作業にはGLM-5.2を選びましょう。曖昧なリポジトリ変更、視覚的なフロントエンドのデバッグ、そして最初の試行に失敗した場合のコストがモデル呼び出しのコストを上回るタスクにはClaude Opus 5を選びます。 より強い主張には注意が必要です。Z.aiは2026年6月にGLM-5.2をリリースしましたが、AnthropicがOpus 5をリリースしたのは7月24日です。コミュニティでの議論や「実環境」の比較の多くは、現在もGLM-5.2とOpus 4.8を比較しています。これらの結果は有用な背景情報ですが、GLM-5.2がOpus 5に勝つ、あるいは負けることの証拠ではありません。 この記事は、 firsthand benchmarkではなく、証拠に基づく比較です。結論は、現行のモデルドキュメント、GPTProtoの料金、独立したベンチマークデータ、ベンダーの開示情報、そしてコミュニティによる評価手法に基づいています。GLM-5.2とOpus 5を直接比較した証拠がまだない場合は、その制限を明示しています。

Michael Johnson | 2026-08-04

コーディングにおけるGLM-5.2 vs Kimi K3:2026年、開発者にとって優れているのはどちら?

コーディングにおけるGLM-5.2 vs Kimi K3:2026年、開発者にとって優れているのはどちら?

TL;DR: 難しいタスク、長時間にわたるタスク、またはビジュアル要素を含むタスクでは、Kimi K3のほうが優れたコーディングモデルです。Moonshotが公開したコーディング比較ではGLM-5.2を上回り、ホスト型サービスを通じて画像と動画も扱えます。一方、日常的なリポジトリ作業のデフォルトとしては、GLM-5.2のほうが適しています。コストが大幅に安く、運用するモデルも小さく、寛容なMITライセンスを採用しているためです。Kimi K3も現在は重みが公開されていますが、1.56 TBのリポジトリ、64基以上のアクセラレータを推奨する構成、独自ライセンスにより、セルフホスティングへの取り組みは大幅に大きくなります。ボトルネックが能力ならKimiを、毎日のコストと運用の簡便さを重視するならGLMを選びましょう。 GLM-5.2とKimi K3 Codeの比較で興味深いのは、どちらもReactコンポーネントを作成したり、短いアルゴリズムを解いたりできることではありません。このレベルのモデルなら、その基準はすでにクリアしています。重要なのは、課題が複雑になったときにどうなるかです。リポジトリの監査、複数ファイルにまたがる移行、スクリーンショットでしか現れないバグ、あるいは複数のシステムの整合性を保つ必要がある、プレイ可能なThree.jsプロトタイプなどです。 価格差が重要になり始めるのも、まさにこの領域です。最も難しい公開テストではKimi K3のほうが優れていますが、公式の出力価格はGLM-5.2の3倍以上です。日常的なレビューを何千件も処理するチームなら、1ドルあたりの処理量ではGLMのほうが多くなる可能性があります。難しいビジュアルプロジェクトを1件救いたい開発者なら、K3のために喜んで料金を支払うでしょう。

Tiffany Layne | 2026-07-28