1M Context and 128K Output
Process long documents, extended conversations, and large retrieval results within a 1M-token context window. Standard requests can return up to 128K tokens, including thinking tokens.
curl --request POST "https://gptproto.com/v1/chat/completions" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "claude-haiku-5-5",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}'Chat, agents de codage et travail documentaire. Facturé par million de tokens — l'entrée, l'entrée en cache et la sortie sont facturées séparément. GPTProto est 10 % en dessous des tarifs officiels.
| Scénario | Liste Claude | OpenRouter | GPTProto | Économies / mois |
|---|---|---|---|---|
| Personnel10M tokens / mois (4.8M en cache) | $1.37 | $1.44 | $1.23 | −$0.14≈ $1.64 / an |
| Équipe100M tokens / mois (48M en cache) | $13.68 | $14.43 | $12.31 | −$1.37≈ $16.42 / an |
| Entreprise500M tokens / mois (240M en cache) | $68.40 | $72.16 | $61.56 | −$6.84≈ $82.08 / an |
Use Anthropic's fast, small model for frequent developer tasks that need low latency, adjustable reasoning, or long input. Claude Haiku 5.5 combines a 1M-token context window, up to 128K output, text and image input, and adaptive thinking controlled through effort.
1M Context and 128K Output
Process long documents, extended conversations, and large retrieval results within a 1M-token context window. Standard requests can return up to 128K tokens, including thinking tokens.
Adaptive Thinking with Effort
Thinking is adaptive and enabled by default. Use the effort setting to balance reasoning depth, response time, and token consumption; medium is the default on Anthropic's API.
Built for High-Volume Steps
Use Haiku 5.5 for classification, extraction, routing, summaries, database queries, support replies, and focused subagent tasks that run frequently or sit inside a larger workflow.
Text and Image Understanding
Send text or images and receive text output. Use it for document questions, screenshot analysis, chart reading, and visual triage; verify image-input support on the selected GPTProto route.
Claude Haiku 5.5 is Anthropic's October 7, 2026 update to its small-model line. Anthropic positions it for latency-sensitive and cost-sensitive work such as classification, extraction, routing, live support, browser use, and focused subagent steps. It accepts text and images and returns text.
The model expands the Haiku 4.5 context window from 200K to 1M tokens and its maximum output from 64K to 128K. It is also the first Haiku model with adaptive thinking and adjustable effort. Thinking is on by default, and thinking tokens count toward max_tokens.
Anthropic's direct API uses prompt-length pricing. Requests with up to 100K input tokens use the lower tier; requests over 100K use a higher rate for the full request. This threshold matters when estimating a Claude Haiku 5.5 API cost for long documents or long-running agents. The live GPTProto pricing module is the source of truth for GPTProto billing.
| Specification | Claude Haiku 5.5 |
|---|---|
| Provider / release | Anthropic / October 7, 2026 |
| Official Anthropic model ID | claude-haiku-5-5 |
| GPTProto model string | claude-haiku-5-5 |
| Input → output | Text and images → text |
| Context window / maximum output | 1M tokens / 128K tokens |
| Thinking / default effort | Adaptive / medium |
| Official direct rate, prompts up to 100K | $0.10 input / $0.50 output per 1M tokens |
| Official direct rate, prompts over 100K | $0.50 input / $2.50 output per 1M tokens |
| Reliable knowledge cutoff / training cutoff | June 2026 / June 2026 |
| Open weights | No; API access does not include downloadable weights |
Classification, extraction, and routing. Label requests, extract fields, route tickets, or choose which model handles the next step. Keep schemas and acceptance checks in the application.
Customer support and live interfaces. Generate short answers, case summaries, query rewrites, and browser assistance. Keep retrieval, permissions, and escalation rules outside the model.
Coding subagents. Assign narrow tasks such as file discovery, change summaries, error triage, or small patches. Use Sonnet 5.5 or Opus 5.5 for more complex, open-ended coding work.
Document and image analysis. Ask questions across long documents or extract data from screenshots and charts. The larger window adds capacity, but citations and validation still matter.
Long tasks and agentic work. Increase effort when a focused step needs more reasoning. Long prompts can cross the 100K pricing threshold, while the newer tokenizer counts the same text as approximately 30% more tokens than Haiku 4.5.
Choose by workload rather than family name alone. Haiku 5.5 is the starting point for frequent, focused steps; Sonnet 5.5 remains the stronger option for difficult coding and longer chains of judgment.
| Decision factor | Haiku 4.5 | Haiku 5.5 | Sonnet 5.5 |
|---|---|---|---|
| Context / maximum output | 200K / 64K | 1M / 128K | 1M / 128K |
| Thinking control | Manual thinking budget | Adaptive; default medium |
Adaptive; default high |
| Official direct input / output rate | $1 / $5 per 1M | $0.10 / $0.50 up to 100K input; $0.50 / $2.50 over 100K | $2 / $10 per 1M |
| OSWorld 2.1, offline subset | 15.7% | 72.4% | 83.9% |
| Terminal-Bench 4.0 | 0.0% | 39.2% | 70.6% |
| Best starting point | Existing Haiku 4.5 workloads not yet migrated | High-volume routing, extraction, support, and subagents | Complex coding and multi-step agent work |
The benchmark figures above are reported by Anthropic and were not independently reproduced by GPTProto. Run the same prompts, tools, time budget, and acceptance checks on your own workload before moving production traffic.
For GPT 6.1 Sol vs Claude Haiku 5.5, start with Haiku for repeated, well-scoped steps and evaluate GPT 6.1 Sol for demanding coding or agent tasks in an OpenAI-style workflow. Do not compare base token prices alone: include retries, thinking tokens, the 100K Haiku pricing threshold, latency, and accepted-task rate.
Moving from Haiku 4.5 requires more than replacing the model name. Test these changes against the exact endpoint shown in GPTProto's fixed API Usage section:
Confirm the model string. Anthropic uses claude-haiku-5-5 with no date suffix; use the exact ID displayed by GPTProto.
Recount tokens. The same text can produce approximately 30% more tokens than on Haiku 4.5. Recalculate budgets and alerts.
Replace manual thinking budgets. budget_tokens can return a 400 error. Use adaptive thinking and control depth through effort.
Parse blocks by type. A response may begin with a thinking block rather than user-facing text.
Remove custom sampling values. Non-default temperature, top_p, or top_k settings can return a 400 error.
Remove assistant prefills. End messages with a user turn.
Handle new failure states. Add logic for refusal and model_context_window_exceeded, then test before changing the default.
If you are migrating from Anthropic to GPTProto, also verify authentication, endpoint compatibility, streaming events, tool calls, image input, and error objects. Native Anthropic parameters do not necessarily map one-to-one onto every OpenAI-compatible route.
Choose Haiku 5.5 for many short or bounded tasks: classification, extraction, routing, retrieval compression, support assistance, visual triage, or focused subagent work. It is also a candidate when Haiku 4.5 lacks enough context or reasoning control.
Choose Claude Sonnet 5.5 for deeper coding judgment or sustained multi-step execution. Keep Claude Haiku 4.5 available until token counts, response parsing, tools, and cost behavior are validated.
Guides, comparaisons et mises à jour liés à ce modèle.
Tous les articles
A developer-focused comparison of Claude Opus 5.5 and GPT-6 Sol for AI coding assistants and autonomous agents. Explore API pricing, prompt caching, context windows, tool calling, and error recovery—plus what to test before choosing a model for your GPTProto coding workflow.

Compare Claude Opus 5.5 and Sonnet 5.5 for coding and agents. See benchmark tradeoffs, cost per task, and GPTProto's 10% lower input/output rates.

Sonnet 5.5 keeps Sonnet 5’s official token rates. Explore coding results, Claude Code access, upgrade checks, and GPTProto’s 10% off API pricing.