Claude Haiku 5.5: a 1M-context small model at $0.10 per million input tokens
In brief
On October 7, 2026 Anthropic released Claude Haiku 5.5, its cheapest and fastest model, with an effort setting and a 1M-token context window. What it costs, where it fits next to Sonnet 5.5, and the API calls that now return 400 errors when you migrate from Haiku 4.5.
Contents
On October 7, 2026, Anthropic released Claude Haiku 5.5 (claude-haiku-5-5). It is the smallest and cheapest model in the Claude 5.5 line, with a 1M-token context window and up to 128k tokens of output. It is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry.
A "small model" in this context means a model that answers faster and costs less per token than the flagship models, and is built for work you run thousands of times: classifying tickets, summarizing documents, turning questions into database queries, compacting long conversations.
What it costs
Prices are per million tokens. Haiku 5.5 has two price tiers, one for requests up to 100k tokens and one above.
| Haiku 5.5 (up to 100k / over 100k) | Haiku 4.5 | Sonnet 5.5 | |
|---|---|---|---|
| Input | $0.10 / $0.50 | $1.00 | $2.00 |
| Output | $0.50 / $2.50 | $5.00 | $10.00 |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
Anthropic says the average workload costs about 75% less than on Haiku 4.5. Haiku 5.5 uses a new tokenizer that produces slightly more tokens for the same text, so measure your own bill instead of dividing the list prices.
Where it fits
Anthropic's benchmarks put Haiku 5.5 well ahead of Haiku 4.5 and behind Sonnet 5.5. A few numbers from the announcement:
- OSWorld 2.1 offline subset (computer use): 72.4% for Haiku 5.5, 15.7% for Haiku 4.5, 83.9% for Sonnet 5.5.
- Terminal-Bench 4.0 (command-line tasks): 39.2% for Haiku 5.5, 70.6% for Sonnet 5.5.
- Humanity's Last Exam without tools: 45.9% for Haiku 5.5, 56.9% for Sonnet 5.5.
Anthropic says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. Haiku 5.5 is aimed at high-volume steps and at acting as a subagent, a helper model that a larger model hands narrow jobs to. Choosing the right Claude model covers how to split work between tiers.
Haiku 5.5 is the first Haiku with the effort parameter (Low, Medium, High, Xhigh, Max). Effort controls how much thinking the model does before answering, so you can pay for more reasoning on hard requests only. Prompt antipatterns and effort calibration explains how to run a sweep to find the right setting.
Migrating from Haiku 4.5
Several request patterns that worked on Haiku 4.5 now return a 400 error on Haiku 5.5:
- A manual
budget_tokensvalue in thethinkingblock. Thinking is adaptive and on by default. temperature,top_p, andtop_k.- Assistant prefill, where you end the messages array with a partial assistant turn.
- The
computer_20250124tool version.
Thinking text is also omitted from responses unless you ask for it:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=2048,
thinking={"type": "adaptive", "display": "summarized"},
output_config={"effort": "low"},
messages=[{"role": "user", "content": "Classify this ticket: 'I was charged twice for May.'"}],
)
Run your test set on Haiku 5.5 with the old parameters removed, then compare quality and cost against Haiku 4.5 before you switch production traffic.
Safeguards
Cyber safeguards on Haiku 5.5 are stricter than on Haiku 4.5 but looser than on Sonnet 5.5. Penetration testing and attacker-oriented techniques remain blocked. Biology safeguards match Sonnet 5.5 and Opus 5.5. Verification programs, such as the Cyber Verification Program, cover wider access.
Other changes announced the same day
- Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens. Anthropic estimates this cuts the cost of most agentic tasks by about 20%, because agents re-read the same context on every step. See prompt caching implementation for how to keep your cache hit rate high.
- Max and Team plans now include monthly API credits: $100 on Max 5x, $200 on Max 20x, and up to $500 pooled for Team.
- The Python and TypeScript SDKs added beta classes for computer use and browser use. You subclass one class and implement one method per tool; the SDK runs the tool loop, the policies, and the approval callback. See computer use and browser use for the tools themselves.
What to do this week
- List the calls in your product that use Haiku 4.5 or a larger model for classification, extraction, or summarizing.
- Re-run a sample of 200 real inputs per call on
claude-haiku-5-5, with the removed parameters stripped out. - Move the calls where quality holds, and record the new cost per call.
Check the release notes for the current migration guidance, since this model is new.