Claude Haiku 5.5 Cuts Token Prices 90%, but Only Under 100,000 Tokens
Anthropic shipped Claude Haiku 5.5 on October 7. Input runs $0.10 per million tokens, a tenth of Haiku 4.5, but any prompt over 100,000 tokens pays $0.50, five times the headline rate.
- Terminal-Bench 4.0: 39.2% for Haiku 5.5, 0.0% for Haiku 4.5.
- Input is
$0.10per million tokens, but a prompt over 100,000 tokens pays $0.50. - The lifetime guarantee on Haiku 4.5 expires October 15.
Anthropic released Claude Haiku 5.5 (claude-haiku-5-5) on October 7, 2026. Haiku is the smallest and cheapest tier in the Claude lineup. It takes the work that comes in volume but does not need deep deliberation: shortening long text, sorting incoming documents into categories, pulling one field out of a filing.
Anthropic's official post calls it "the cheapest, fastest, and most capable small model we've ever released." Input is $0.10 per million tokens and output is $0.50, down from $1.00 and $5.00 on Haiku 4.5, a tenth of the previous rate. Those rates apply only to requests whose prompt stays under 100,000 tokens. Above that line they become $0.50 and $2.50, five times higher.
39.2% on a test where Haiku 4.5 scored 0.0%
Set the price aside for a moment and look at what the model does that its predecessor could not. In the official table on the announcement page, the widest gap is in agentic coding.
Terminal-Bench 4.0 measures whether a model can drive a long task in a terminal to completion on its own. Haiku 5.5 scores 39.2%. Haiku 4.5 scored 0.0% on the same test. A zero there is not a failed measurement; it means not a single task finished. Until now the Haiku tier was not a candidate for coding agents at all, and this release is the first time it is. OpenAI's small model GPT-6 Luna scores 16.4% on the same test.
The gap in driving a computer directly is nearly as wide. Haiku 5.5 takes 72.4% on the OSWorld 2.1 offline partial score, against 15.7% for Haiku 4.5 and 48.9% for GPT-6 Luna. One tier up, Sonnet 5.5 posts 83.9% there and 70.6% on Terminal-Bench. The announcement page says outright that Sonnet 5.5 and Opus 5.5 remain the better choice for agentic coding at Terminal-Bench complexity. What Haiku 5.5 replaces is not a larger model. It replaces Haiku 4.5.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| OSWorld 2.1 (offline) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | not published | 64.5% |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
The Haiku tier also gets an effort dial for the first time. Effort controls how long the model deliberates before it answers, across five settings from low to max, and Haiku 5.5 defaults to medium. Sonnet 5.5 has a band where pushing effort to max drives cost per task above the larger models, and that same arithmetic now applies to Haiku.
One independent measurement is in. Artificial Analysis records an intelligence index of 38 at high effort, 10th out of 182 models, with output at 136.6 tokens per second against a comparison median of 110.9. Time to first token, though, is 26.15 seconds, where the median for comparable reasoning models is 2.17 seconds. The "fastest" in Anthropic's claim is the rate at which tokens come out. The wait before the first character appears grows with the effort setting. Anywhere perceived latency matters, such as live customer support, the setting has to come down and the measurement has to be redone.
Cross 100,000 tokens and the rate goes up 5x
| Per 1M tokens | Haiku 5.5 (under 100k) | Haiku 5.5 (over 100k) | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| Batch input | $0.05 | $0.25 | $0.50 |
Haiku 5.5 is the only model in the current Claude lineup priced by prompt length. The pricing docs state that "Claude 4.6 and later models (except Claude Haiku 5.5) ... include the full 1M token context window at standard pricing," with the parenthetical that "a 900k-token request is billed at the same per-token rate as a 9k-token request." Fable 5.1, Opus 5.5, and Sonnet 5.5 hold one rate no matter how long the prompt gets. Haiku 5.5 alone has a step at 100,000 tokens.
That line sits closer than the number suggests. Haiku 5.5 uses the newer tokenizer shared with Sonnet 5.5 and Opus 5.5, and the pricing docs say that tokenizer produces roughly 30% more tokens for the same text. Haiku 4.5 uses the previous one. So a document that counted 77,000 tokens on Haiku 4.5 counts over 100,000 on Haiku 5.5.
One document, same length
Recompute the discount as cost to process one document rather than cost per token, and it moves. In the short band, $1.00 becomes $0.13 ($0.10 times 1.30 more tokens), an 87% cut. In the long band it only falls to $0.65, a cut of 35%. That is why Anthropic's marketing number of 90% and its own footnote figure of "around 75% on average" disagree. The footnote states that its number accounts for the token increase.
The uses Anthropic recommends happen to straddle that line. The announcement page lists summarization, classification, database queries, and context compaction as the work Haiku 5.5 is for. Compaction summarizes the earlier part of a conversation to free room once it outgrows the model's window, which is what the Claude API ships as on-demand compaction. The input to a compaction call is the long conversation being compacted, so that use lands in the long band by construction. A subagent reading several repository files crosses 100,000 tokens quickly too. The work that actually gets the tenth-of-the-price rate is the small-prompt kind: routing a support ticket, extracting one number from a filing.
The announcement page says about 90% of Haiku 4.5 requests fell in the short band. Whether your own workload sits inside that 90% is a question the token distribution answers, not the average.
Sonnet 5.5 changed the same day. Cache reads went from $0.20 to $0.10 per million tokens, a 50% cut. Teams that park a long system prompt in the cache get that reduction without switching models.
Can you use it today
| Item | Detail |
|---|---|
| Who it is for | Developers with an API key, and Claude Code users |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Model ID | claude-haiku-5-5 (anthropic.claude-haiku-5-5 on Bedrock) |
| Price | From $0.10 per 1M input tokens, $0.50 above 100,000 tokens |
| Requirements | None. No waitlist, beta application, admin approval, or verification program |
| Regional availability | Claude API is open everywhere. Bedrock geo inference profiles are US, EU, AU, JP, and global, with no APAC profile |
An API key is all it takes. The Claude API is open regardless of where you call from, the context window is 1M tokens, maximum output is 128,000 tokens, and the reliable knowledge cutoff is June 2026.
If you have to pin inference to a named geography, the answer changes. According to the AWS machine learning blog, Bedrock offers five geographic inference profiles for Haiku 5.5: US (us.), EU (eu.), AU (au.), JP (jp.), and global (global.). There is no APAC profile, so a caller in Singapore or the rest of Southeast Asia has JP and AU as the nearest named options and global. as the default. In-region inference, where the request never leaves one Region, is a separate path that tops out at the 5 generation in Singapore (ap-southeast-1) and Seoul, so Haiku 5.5 is not on it either.
For most APAC teams that is a contract question rather than a legal one. Singapore's PDPA does not require personal data to stay on the island, only that an overseas recipient be bound to a comparable standard, and MAS imposes no residency rule on financial institutions. But a customer agreement, a procurement questionnaire, or an internal policy that already named a processing location will not accept global., and for those Haiku 5.5 has no answer today. On the Claude API, the inference_geo: "us" option that pins inference to the United States adds 1.1x to every billed item.
Teams still running the previous model have a date on the calendar. The deprecation docs still list claude-haiku-4-5-20251001 as Active, with an estimated retirement of "not sooner than October 15, 2026." That date is not a shutdown. It is the end of the period Anthropic committed to keeping the model alive. After it passes, a deprecation notice can arrive at any time, and once it does the model retires at least 60 days later. The most recent precedent is Sonnet 4.5: notice on September 30, retirement set for November 30.
If you run a pipeline on Haiku 4.5 today, count which band your prompts fall into before you change the model ID. The token counting endpoint (POST /v1/messages/count_tokens) counts using the tokenizer of whichever model you pass in model, and the call itself is free. Send a real sample of your prompts through claude-haiku-4-5-20251001 and claude-haiku-5-5 and compare the two input_tokens values. Anthropic's own migration docs recommend this. Once you know what share of requests exceeds 100,000 tokens, you know whether your monthly bill drops 87% or 35%, and moving with that number in hand means you will not be splitting 60 days once the deprecation notice lands after October 15.