Anthropic Releases Claude Haiku 5.5 With a 1M-Token Context Window and Input Pricing From $0.10 per Million Tokens, Down From $1 on Haiku 4.5
Anthropic's Claude Haiku 5.5 lists at $0.10 per million input tokens for prompts up to 100,000 tokens, versus $1 for Haiku 4.5, and adds a 1M-token context window and adaptive thinking.
Editor's Note ·
- Correction:
- The article says "Per-token prices understate the change in cost for existing workloads" and that the tokenizer is "why Anthropic's own average-cost estimate is smaller than the per-token price cut." The roughly 30% more tokens per text mean per-token prices overstate the savings, not understate them. Anthropic's announcement (footnote 2) says the 75% average figure reflects both the price cuts (90% for requests up to 100,000 tokens, 50% above that, with about 90% of Haiku 4.5 requests in the lower bucket) and the tokenizer change, not the tokenizer alone.
- Clarification:
- The article gives OSWorld 2.1 scores without qualification. The announcement labels those figures "Offline subset", and its Sonnet 5.5 FrontierCode 1.1 (Main) figure of 52.1% is marked as Xhigh effort.
Overview
Anthropic released Claude Haiku 5.5 on October 7, 2026. According to the Claude Platform release notes, the model (API ID claude-haiku-5-5) is tuned for high-volume and latency-sensitive work, has a 1M token context window and 128k maximum output tokens, and supports adaptive thinking with the effort parameter. Anthropic’s announcement calls it the “cheapest, fastest, and most capable small model we’ve ever released.”
What We Know
Pricing
Anthropic’s pricing page lists Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. For prompts over 100,000 tokens, the page lists $0.50 per million input tokens and $2.50 per million output tokens. Claude Haiku 4.5 is listed at $1 per million input tokens and $5 per million output tokens.
The tiered structure is a change from other recent models: the same page says Claude 4.6 and later models, except Haiku 5.5, include the full 1M token context window at standard pricing, while Haiku 5.5 is priced by prompt length.
Per-token prices understate the change in cost for existing workloads. Anthropic’s documentation for the release says Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Haiku 4.5. The announcement says that on average the model now costs around 75% less to run than Haiku 4.5.
Capabilities and limits
The documentation says the context window grows to 1M tokens and maximum output to 128k tokens, up from 200k and 64k on Haiku 4.5. Anthropic’s announcement positions the model for summaries, compactions, database queries and classification, as a subagent alongside Opus 5.5 and Sonnet 5.5 on coding work, and for speed-sensitive uses such as live customer support and browser use.
Benchmarks
Anthropic’s announcement includes a comparison table against Haiku 4.5, Sonnet 5.5 and OpenAI’s GPT-6 Luna. These are vendor-reported figures:
- Terminal-Bench 4.0: Haiku 5.5 scores 39.2%, against 0.0% for Haiku 4.5, 16.4% for GPT-6 Luna and 70.6% for Sonnet 5.5.
- OSWorld 2.1: 72.4% for Haiku 5.5, against 15.7% for Haiku 4.5, 48.9% for GPT-6 Luna and 83.9% for Sonnet 5.5.
- FrontierCode 1.1 (Main): 46.4% for Haiku 5.5, 42.4% for GPT-6 Luna and 52.1% for Sonnet 5.5.
- Humanity’s Last Exam without tools: 45.9% for Haiku 5.5, against 10.2% for Haiku 4.5 and 56.9% for Sonnet 5.5.
On each of these rows Sonnet 5.5 scores higher than Haiku 5.5, while Haiku 5.5 scores above GPT-6 Luna on the rows where the table lists both.
Availability
The release notes list the model as available on the Claude API, Claude in Amazon Bedrock, Claude Platform on AWS, Claude on Google Cloud and Claude in Microsoft Foundry. GitHub said the same day that the model is generally available in GitHub Copilot, with a gradual rollout, and that in early testing it matched Claude Sonnet 5 on many coding tasks while using significantly fewer tokens and steps.
Migration changes
The documentation lists several breaking changes from Haiku 4.5. Manual extended thinking with budget_tokens returns an error, non-default sampling parameters (temperature, top_p, top_k) return an error, and assistant message prefill returns an error. Because adaptive thinking is on by default, a response can begin with thinking blocks, so code should select content blocks by type rather than position. Thinking tokens also count toward max_tokens, so a small limit can stop output after a thinking block and before any text.
What We Don’t Know
- The benchmark figures come from Anthropic’s own announcement. The sources reviewed for this article do not include independent evaluations of Haiku 5.5.
- The announcement’s 75% average cost reduction is Anthropic’s estimate; actual savings depend on each workload’s prompt lengths, output lengths and thinking effort.
- The sources reviewed do not state how the 100,000-token pricing threshold affects workloads that mix short and long prompts.
Analysis
The release pairs a large cut in list prices with a tokenizer that produces roughly 30% more tokens for the same text, which is why Anthropic’s own average-cost estimate is smaller than the per-token price cut. Teams running Haiku 4.5 at scale will need to recount prompts and revisit max_tokens settings, as the documentation advises, before comparing bills.