Long-Horizon Agentic Work
Handles multi-stage tasks that run for hours, use multiple tools, recover from failed steps, and continue toward a defined result with less supervision.
Estimate a request with real work scenarios. GPTProto token pricing is 10% below official rates.
Top-up $100 and you get:
Top-up credits with permanent validity. You will receive a total of $100.00.
Additional 10% model discount, saving $11.0937 versus direct official Claude API calls.
Build coding agents, research systems, and complex document workflows with Claude Fable 5.1 API. It combines a 1M-token context window, up to 128K output tokens, and always-on adaptive thinking. Access it on GPTProto at 10% below Anthropic’s standard input and output rates.
Handles multi-stage tasks that run for hours, use multiple tools, recover from failed steps, and continue toward a defined result with less supervision.
Supports repository-wide features, code review, performance analysis, test creation, root-cause debugging, and implementation across multiple files.
Process large repositories and technical documentation within a 1M-token context window, then return up to 128K tokens in one response.
Reasoning stays enabled. Tune effort from low to max to balance task difficulty, latency, and token use instead of switching thinking off.
Claude Fable 5.1 is Anthropic’s most capable generally available model for demanding reasoning and long-horizon agentic work. Released on September 1, 2026, it succeeds Claude Fable 5 while retaining the same 1M-token context window, 128K maximum output, and base list price. It is a proprietary hosted model designed for difficult work that must maintain quality across many steps.
| Specification | Claude Fable 5.1 |
|---|---|
| Release date | September 1, 2026 |
| Official API model ID | claude-fable-5-1 |
| Context window | 1,000,000 tokens |
| Maximum output | 128,000 tokens |
| Thinking | Adaptive, always on |
| Effort | High by default; low to max supported |
| Knowledge cutoff | June 2026 |
Fable 5.1 improves most clearly on work combining reasoning, tools, and long execution chains. Anthropic reports 55.8% on Terminal-Bench 4.0 versus 42.0% for Fable 5. Its direct cache-read price also fell from $1 to $0.25 per million tokens, with Anthropic estimating about 25% lower typical workload cost and savings of up to approximately 45% for highly agentic work. Actual savings depend on cache reuse.
| Application | Why Fable 5.1 fits |
|---|---|
| Agentic coding | Plans repository changes, writes tests, checks results, and recovers from failed steps. |
| Long-running agents | Maintains direction across extended tool sequences and asynchronous work. |
| Research | Synthesizes large source collections and follows multi-part instructions. |
| Knowledge work | Produces document, spreadsheet, presentation, and operational deliverables for review. |
Anthropic’s published benchmark set indicates that Fable 5.1 is intended for tasks where long-horizon completion quality matters more than minimum latency or token price.
| Anthropic-published evaluation | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 55.8% | 42.0% | 52.3% | 37.3% |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
These are provider-run results, not independent rankings. For GPT-5.6, Grok 4.6, Qwen3.8-Max-0902, Kimi K3, or GLM-5.3 Flash comparisons, use identical tools, context, time limits, and acceptance tests. Choose Fable 5.1 when failed completion across a long workflow is more expensive than its higher token price.
Moving from Claude Fable 5 or Opus 5 requires more than changing the model name:
Forced tool choice is unsupported. tool_choice values any and tool return a 400 error. Use auto with explicit instructions and strict tool schemas, or request structured JSON.
Thinking cannot be disabled. Control spend and latency through effort; test medium for routine stages and reserve xhigh or max for evaluation-proven gains.
Keep conversation history append-only. Editing earlier messages can invalidate preserved thinking blocks. After client-side compaction, remove stale thinking blocks or restart from a clean summary.
Use Claude Fable 5.1 for work that is difficult, long-running, and expensive to restart: repository-scale engineering, open-ended debugging, multi-document research, or agents coordinating several tools. Use a lower-cost model for routine chat, extraction, classification, and short code generation.
Guides, comparisons, and updates related to this model.
All Articles
Claude Fable 5.1 and Mythos 5.1 share one model but differ in access. See pricing, features, benchmarks, safeguards, and API migration changes.

Meta Description: Compare GLM 5.3 Flash vs DeepSeek V4 Flash for coding, frontend work, agents, speed, context, and API pricing to choose the better model.

Qwen3.8-Flash-Next vs GLM-5.3 Flash compared for coding, agents, frontend work, speed, pricing, context, licenses, and production use.

What is Tencent Hy4 Preview? See its release status, 770B MoE design, 1M context, API pricing, benchmarks, limits, and model comparisons.