Anthropic has launched Claude Haiku 5.5, the first update to its smallest model in nearly a year, with API prices cut sharply from Haiku 4.5. Per The New Stack, it’s Anthropic’s third 5.5 model in a month, following Opus 5.5 and Sonnet 5.5.
Anthropic calls it its “cheapest, fastest, and most capable small model” yet, in a post on X.
Pricing
Haiku 4.5 cost $1 per million input tokens and $5 per million output tokens, whatever the request size. Haiku 5.5 charges $0.10 and $0.50 for requests under 100,000 tokens. Larger requests cost $0.50 and $2.50. Those are 90% and 50% cuts respectively.
Anthropic says about 90% of Haiku 4.5 requests fell into the cheaper bracket. It puts the average saving at around 75%, a figure that allows for a request mix and an updated tokenizer that uses slightly more tokens per task.

Haiku 5.5 is also the first Haiku with effort controls. The default is medium, and developers can use the setting to cap how many tokens the model spends on a task.
Benchmarks
These are Anthropic’s own figures, as reported by The New Stack. It compares the model with Haiku 4.5 and OpenAI’s GPT-6 Luna, which launched alongside GPT-6 Sol last month.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| OSWorld 2.1 (offline subset, computer use) | 72.4% | 15.7% | 48.9% | Ahead of Haiku 5.5 |
| GDPval-AA v2.1 (knowledge work) | 1,620 | 735 | 1,437 | Ahead of Haiku 5.5 |
| Terminal-Bench 4.0 (command line) | 39.2% | 0% | 16.4% | 70.6% |
Anthropic doesn’t compare the model with Chinese rivals. On Artificial Analysis’ GDPval-AA v2.1, Z.ai’s GLM-5.3-Flash scores 1,647, above Haiku 5.5’s 1,620. Alibaba’s Qwen3.7 Flash is cheaper still, at $0.03 and $0.13 per million input and output tokens for inputs up to 32,000 tokens.
Use cases and availability
Small models have mostly handled high-volume work such as summarisation, classification and routing. Anthropic also pitches Haiku 5.5 for compaction, database queries and speed-sensitive agent work like live customer support and browser use. It’s adding beta computer-use and browser-use support to its Python and TypeScript SDKs.
Haiku 5.5 is available on the Claude Platform (as `claude-haiku-5-5`), Amazon Web Services, Google Cloud and Microsoft Azure. Its cybersecurity safeguards are tighter than Haiku 4.5’s, though Anthropic says they allow more defensive work than Sonnet 5.5’s. They still block penetration testing. Organisations that need broader cyber or biology access can apply to Anthropic’s verification programmes.
Other changes
Anthropic is halving Sonnet 5.5’s cache-read price, from $0.20 to $0.10 per million tokens. It says that makes most agentic work about 20% cheaper. The cut rolls out across platforms from Wednesday, though some existing Azure and Google Cloud customers will wait a few days.
Max and Team subscribers also get monthly API credits this week, usable with any model on the Claude Platform. Max 5x gets $100 a month and Max 20x gets $200. Team plans get up to $500, pooled across users.
Does Haiku 5.5 replace Haiku 4.5?
It’s the new Haiku, but the sources don’t say when Haiku 4.5 retires. Check Anthropic’s deprecation notices before migrating production workloads.
Is Haiku 5.5 cheaper for every request?
Yes, but by different margins. Requests under 100,000 tokens cost 90% less, while larger ones cost 50% less. The tokenizer change means you may use slightly more tokens for the same task.
Can I use it in the UAE?
The sources don’t give a regional breakdown. It’s offered through the Claude Platform and the three major clouds, so check availability in your account.


















