◈AI Codex
Models & Pricingupdate

Claude Haiku 5.5: a 1M-context small model at $0.10 per million input tokens

In brief

On October 7, 2026 Anthropic released Claude Haiku 5.5, its cheapest and fastest model, with an effort setting and a 1M-token context window. What it costs, where it fits next to Sonnet 5.5, and the API calls that now return 400 errors when you migrate from Haiku 4.5.

6 min read·Context Window

Contents

♡Sign in to save

On October 7, 2026, Anthropic released Claude Haiku 5.5 (claude-haiku-5-5). It is the smallest and cheapest model in the Claude 5.5 line, with a 1M-token context window and up to 128k tokens of output. It is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry.

A "small model" in this context means a model that answers faster and costs less per token than the flagship models, and is built for work you run thousands of times: classifying tickets, summarizing documents, turning questions into database queries, compacting long conversations.

What it costs

Prices are per million tokens. Haiku 5.5 has two price tiers, one for requests up to 100k tokens and one above.

Haiku 5.5 (up to 100k / over 100k) Haiku 4.5 Sonnet 5.5
Input $0.10 / $0.50 $1.00 $2.00
Output $0.50 / $2.50 $5.00 $10.00
Cache reads $0.01 / $0.05 $0.10 $0.10
Cache writes $0.125 / $0.625 $1.25 $2.50

Anthropic says the average workload costs about 75% less than on Haiku 4.5. Haiku 5.5 uses a new tokenizer that produces slightly more tokens for the same text, so measure your own bill instead of dividing the list prices.

Where it fits

Anthropic's benchmarks put Haiku 5.5 well ahead of Haiku 4.5 and behind Sonnet 5.5. A few numbers from the announcement:

  • OSWorld 2.1 offline subset (computer use): 72.4% for Haiku 5.5, 15.7% for Haiku 4.5, 83.9% for Sonnet 5.5.
  • Terminal-Bench 4.0 (command-line tasks): 39.2% for Haiku 5.5, 70.6% for Sonnet 5.5.
  • Humanity's Last Exam without tools: 45.9% for Haiku 5.5, 56.9% for Sonnet 5.5.

Anthropic says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. Haiku 5.5 is aimed at high-volume steps and at acting as a subagent, a helper model that a larger model hands narrow jobs to. Choosing the right Claude model covers how to split work between tiers.

Haiku 5.5 is the first Haiku with the effort parameter (Low, Medium, High, Xhigh, Max). Effort controls how much thinking the model does before answering, so you can pay for more reasoning on hard requests only. Prompt antipatterns and effort calibration explains how to run a sweep to find the right setting.

Migrating from Haiku 4.5

Several request patterns that worked on Haiku 4.5 now return a 400 error on Haiku 5.5:

  • A manual budget_tokens value in the thinking block. Thinking is adaptive and on by default.
  • temperature, top_p, and top_k.
  • Assistant prefill, where you end the messages array with a partial assistant turn.
  • The computer_20250124 tool version.

Thinking text is also omitted from responses unless you ask for it:

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2048,
    thinking={"type": "adaptive", "display": "summarized"},
    output_config={"effort": "low"},
    messages=[{"role": "user", "content": "Classify this ticket: 'I was charged twice for May.'"}],
)

Run your test set on Haiku 5.5 with the old parameters removed, then compare quality and cost against Haiku 4.5 before you switch production traffic.

Safeguards

Cyber safeguards on Haiku 5.5 are stricter than on Haiku 4.5 but looser than on Sonnet 5.5. Penetration testing and attacker-oriented techniques remain blocked. Biology safeguards match Sonnet 5.5 and Opus 5.5. Verification programs, such as the Cyber Verification Program, cover wider access.

Other changes announced the same day

  • Sonnet 5.5 cache reads dropped from $0.20 to $0.10 per million tokens. Anthropic estimates this cuts the cost of most agentic tasks by about 20%, because agents re-read the same context on every step. See prompt caching implementation for how to keep your cache hit rate high.
  • Max and Team plans now include monthly API credits: $100 on Max 5x, $200 on Max 20x, and up to $500 pooled for Team.
  • The Python and TypeScript SDKs added beta classes for computer use and browser use. You subclass one class and implement one method per tool; the SDK runs the tool loop, the policies, and the approval callback. See computer use and browser use for the tools themselves.

What to do this week

  1. List the calls in your product that use Haiku 4.5 or a larger model for classification, extraction, or summarizing.
  2. Re-run a sample of 200 real inputs per call on claude-haiku-5-5, with the removed parameters stripped out.
  3. Move the calls where quality holds, and record the new cost per call.

Check the release notes for the current migration guidance, since this model is new.

Related tools

Weekly brief

For people actually using Claude at work.

Each week: one thing Claude can do in your work that most people haven't figured out yet — plus the failure modes to avoid. No tutorials. No hype.

No spam. Unsubscribe anytime.

What to read next

Picked for where you are now

All articles →