Use Claude Haiku 5.5 for narrow, high-volume jobs whose output you can check cheaply, such as classification, extraction, summaries, routing and subagent lookups. It costs 20 times less per token than Sonnet 5.5 on prompts up to 100,000 tokens. Keep Sonnet 5.5 (or Opus 5.5) for complex agentic coding and any task where a wrong answer is costly: Sonnet solves 70.6% of Terminal-Bench 4.0 tasks to Haiku’s 39.2%.
Claude Haiku 5.5 launched on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, one tenth of Haiku 4.5’s price. The more useful question is not whether it is cheap, but which jobs it can now take away from Sonnet 5.5.
The short answer from the published numbers: Haiku 5.5 fits narrow, high-volume work such as classification, extraction, summarization, compaction and routing, plus lookups done by subagents under a larger model. Sonnet 5.5 keeps complex agentic coding and any task where a wrong answer is expensive. Anthropic says this itself, and its own benchmark table backs it up.
The price cut is real, but smaller than the headline. Anthropic puts the average saving at about 75%, not 90%, and the discount shrinks to 50% once a prompt passes 100,000 tokens. The rest of this piece shows where those limits sit and how to decide, job by job.
Haiku 5.5 vs Sonnet 5.5 at a glance
Sonnet 5.5 costs 20 times more per input or output token than Haiku 5.5 on prompts up to 100,000 tokens, and in return scores higher on every benchmark Anthropic published.
| Haiku 5.5 | Sonnet 5.5 | |
|---|---|---|
| Input / output per 1M tokens | $0.10 / $0.50 (up to 100K), $0.50 / $2.50 above | $2 / $10 |
| Cache read per 1M tokens | $0.01 (up to 100K), $0.05 above | $0.10 |
| Latency tier | Fastest | Fast |
| Default effort | medium | high |
| Context window / max output | 1M / 128K tokens | 1M / 128K tokens |
| Knowledge cutoff | June 2026 | June 2026 |
| Anthropic’s positioning | High-volume, latency-sensitive work such as classification, extraction and routing | Best combination of speed and intelligence |
What Anthropic announced
Anthropic calls Haiku 5.5 the cheapest, fastest and most capable small model it has released. It first named the model in the Opus 5.5 launch post on September 22 and shipped it on October 7, after Sonnet 5.5 on September 28. It is the first Haiku in about a year, and there was no Haiku 5.
| Spec | Claude Haiku 5.5 |
|---|---|
| API model ID | claude-haiku-5-5 |
| Context window | 1M tokens |
| Max output | 128K tokens (300K in the Batch API beta) |
| Input / output | Text and images in, text out |
| Thinking | Adaptive, default effort medium |
| Knowledge cutoff | June 2026 |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Retirement | Not before October 7, 2027 |
It is also the first Haiku with an adjustable effort setting. Anthropic positions it for summaries, compactions, database queries and classification, as a subagent under Opus 5.5 or Sonnet 5.5 on coding work, and for speed-sensitive tasks such as live customer support and browser use. Its footnote adds that Haiku 5.5 is the fastest model at standard speed, but slower than Opus in Fast Mode.
Two other changes shipped the same day and matter for any cost comparison:
- Sonnet 5.5 cache reads were cut by half, from $0.20 to $0.10 per million tokens. Anthropic says this makes Sonnet 5.5 about 20% cheaper on most agentic work.
- Max and Team subscribers get a monthly API credit for the Claude Platform: $100 for Max 5x, $200 for Max 20x and up to $500 pooled for Team. The Python and TypeScript SDKs also gained beta support for computer use and browser use.
The price headline: 90%, 75% and the 100K cliff
Haiku 5.5 is priced in two tiers by prompt size, and the cheap tier is the one the headline describes.
| Price per 1M tokens | Haiku 5.5 (prompt up to 100K) | Haiku 5.5 (prompt over 100K) | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| Cache write (5 min) | $0.125 | $0.625 | $1.25 |
The Batch API takes a further 50% off input and output. Thinking tokens bill as output, so a higher effort setting costs more per request.
Three details explain why Anthropic says “about 75% cheaper on average” instead of 90%:
- The 100K cliff. Above 100,000 tokens the rates are five times higher, so the saving versus Haiku 4.5 drops from 90% to 50%. Anthropic says about 90% of Haiku 4.5 requests fell below that line, but it does not publish how much of the spend did. A workload built on long documents or growing agent context will save less.
- The tokenizer. Haiku 5.5 uses the newer tokenizer shared with Sonnet 5.5 and Opus 5.5. The platform docs say the same text counts as roughly 30% more tokens than on Haiku 4.5, which Anthropic’s launch footnote describes only as “slightly more”. Your bill rises by that factor before the discount applies.
- Thinking is on by default. Adaptive thinking is enabled unless you turn it off, and thinking tokens bill as output, so the effort level feeds directly into cost.
To see what the tokenizer does, take a request with 2,000 input and 200 output tokens on Haiku 5.5. The same text would be about 1,540 and 155 tokens on Haiku 4.5. At published rates, one million such requests cost $300 on Haiku 5.5 and about $2,300 on Haiku 4.5, a saving near 87% rather than 90%. This is my arithmetic from the published rates and the 30% figure, and it ignores thinking tokens.
The practical takeaway: count your own tokens on the new tokenizer and check what share of your traffic exceeds 100K before you forecast savings.
One comparison point: OpenAI’s GPT-6 Luna lists the same $0.10 and $0.50 base rates, but its surcharge starts above 272,000 input tokens, not 100,000, according to VentureBeat’s check of official pricing.
What the benchmarks say, and what they hide
Haiku 5.5 beats Haiku 4.5 by a wide margin on every benchmark Anthropic published, and leads GPT-6 Luna wherever both appear. Sonnet 5.5 still wins every row. All figures below are vendor-reported, not independently verified.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 | Haiku as share of Sonnet |
|---|---|---|---|---|---|
| FrontierCode 1.1 (main) | 46.4% | not reported | 42.4% | 52.1% (xhigh effort) | 89% |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 48.9% | 83.9% | 86% |
| Humanity’s Last Exam (no tools) | 45.9% | 10.2% | not reported | 56.9% | 81% |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% | 75% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% | 56% |
| GDPval-AA v2.1 (Elo-style score) | 1620 | 735 | 1437 | 1840 | not a ratio |
| AA-Briefcase v1.1 (Elo-style score) | 1578 | 614 | 1336 | 1824 | not a ratio |
The last column is my own calculation from Anthropic’s table. It shows where the gap to Sonnet is narrow and where it is wide. Haiku 5.5 gets within 11% to 14% of Sonnet on FrontierCode and OSWorld, but reaches only 56% of Sonnet’s Terminal-Bench score. That single gap is the clearest sign of where the cheaper model stops being a substitute.
Two caveats change how to read the table:
- Effort setting. VentureBeat reports that Anthropic’s Terminal-Bench chart places the 39.2% score at maximum effort, with medium, the API default, near 20%. Check which effort level a benchmark used before comparing it with your own default-effort results.
- Customer evidence is narrow. Box reports an 11-point gain over Haiku 4.5 at about half the latency, AlphaSense a 0.84 vs 0.76 score on 400 queries for a workload of about 8M calls a week, and HubSpot a 92.8% average on its CRM suite. Asana’s reported 30% latency drop and up to 2.5x faster inference per turn compare against an unnamed model. None of these give a tokens-per-second figure, and none compare across providers.
Anthropic’s own conclusion is consistent with the data: Sonnet 5.5 and Opus 5.5 stay the better choice for complex agentic coding, and Haiku 5.5 suits narrower tasks that were previously too expensive to run.
Which jobs go to Haiku, and which stay on Sonnet
Route by the job, not by the model’s reputation. The table starts from the use cases Anthropic names and the benchmark gaps above; rows marked “my judgment” are mine, not Anthropic’s.
| Job | Route to | Why |
|---|---|---|
| Ticket tagging, routing, moderation queues | Haiku 5.5, low effort | Narrow, high-volume work Anthropic names; classification and routing are the stated design target |
| Question answering over one or a few documents | Haiku 5.5, medium effort | AlphaSense reports about 8M such calls a week and a 0.84 vs 0.76 score over Haiku 4.5 |
| Summaries and context compaction inside agent loops | Haiku 5.5 | Named use case; cache reads cost $0.01 per million tokens at up to 100K |
| Subagent lookups under an Opus or Sonnet orchestrator | Haiku 5.5 | Rogo’s example: a larger model builds the deck, Haiku pulls one revenue line from a 10-K |
| Browser and desktop automation | Haiku 5.5, test first | 72.4% on OSWorld 2.1 vs 83.9% for Sonnet; SDK support is in beta |
| Multi-step coding and terminal agents | Sonnet 5.5 or Opus 5.5 | Terminal-Bench 4.0: 39.2% vs 70.6% |
| High-stakes calls with no cheap check (my judgment) | Sonnet 5.5 or above | A wrong answer costs more than the token savings |
| Penetration testing and offensive security work | Not Haiku by default | Its safeguards block it; broader access goes through the Cyber Verification Program |
The cost-per-success test
Per-token price is not the whole cost, because a cheaper model that fails more often costs more per correct answer. Take Terminal-Bench 4.0 as a stress case. Sonnet 5.5 costs $2 per million input tokens against Haiku 5.5’s $0.10 for prompts up to 100K, a 20x gap. Haiku solves 39.2% of tasks and Sonnet 70.6%.
If each attempt used the same number of tokens, which is a simplifying assumption, one Haiku success would cost about 2.6 Haiku-attempt units and one Sonnet success about 28. Haiku would be roughly 11 times cheaper per success, but only when you can detect a failure and retry. Above 100K tokens the price gap shrinks to 4x and the advantage drops to about 2x.
This points to the question that decides most routing: can you check the output cheaply? Extraction that must match a schema, classification you can sample against labels, and code that passes tests can all be verified, so Haiku plus a retry wins. A legal summary or a financial sign-off cannot be verified without a person, so the extra accuracy of a larger model is cheap insurance.
Four questions before you route a job
- Is the task narrow and well specified?
- Can a script, schema or sample check catch a wrong answer?
- Does the prompt stay under 100,000 tokens?
- Is a wrong answer cheap to fix?
Four yeses mean Haiku 5.5. One no means test it on your own data before moving. Two or more noes mean stay on Sonnet 5.5 or Opus 5.5. Opus 5.5 sits above Sonnet at $4 and $20 per million tokens, and Anthropic names both as better choices than Haiku for complex agentic coding, so the same four questions apply when choosing between Haiku and Opus.
Worked examples at published rates
These figures are my arithmetic from Anthropic’s list prices. They exclude thinking tokens, caching and the Batch API discount, and they assume both models see the same token counts.
| Scenario | Haiku 5.5 | Sonnet 5.5 | Sonnet costs |
|---|---|---|---|
| A: 1M requests, 2,000 input and 200 output tokens each | $300 | $6,000 | 20x more |
| B: 100K requests, 150K input and 1K output tokens each | $7,750 | $31,000 | 4x more |
| C: same job as B, with prompts trimmed to 90K input tokens | $950 | $19,000 | 20x more |
Scenario A is the clean case: the work stays under 100K tokens, so the full 20x price gap applies. Scenario B shows the cliff. The same model costs 8 times more than in scenario C because every token is billed at the higher rate once a prompt crosses 100K. Trimming prompts by 40% cut Haiku’s bill by 88%, but Sonnet’s by only 39%, because its listed price has no such threshold.
That makes prompt size a design choice worth testing: chunking a document or summarizing earlier turns can move a job under the line, provided accuracy holds. Anthropic’s launch materials do not say how a prompt of exactly 100,000 tokens is billed or which tokens count toward the threshold, so leave a margin and confirm with a test request.
The Sonnet price cut matters here too. Halving Sonnet 5.5’s cache-read price to $0.10 narrows the gap on cache-heavy agent loops. At up to 100K tokens, Sonnet’s cache read is now 10 times Haiku’s $0.01, down from 20 times.
Caveats before you switch
Safeguards are tighter for security work. Anthropic says Haiku 5.5’s cybersecurity safeguards are stricter than Haiku 4.5’s but looser than those on other recent models. They allow a wider range of defensive tasks than Sonnet 5.5’s safeguards, yet still block penetration testing and techniques more likely to be used by attackers. Its biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5. Organizations needing broader access can apply to the Cyber Verification Program or the Life Sciences Verification Program. On alignment, Anthropic reports far fewer instances of misaligned behavior and a lower willingness to cooperate with misuse than Haiku 4.5, with details in the system card.
Swapping the model ID is not enough. The docs list several changes from Haiku 4.5:
- Manual thinking budgets (
budget_tokens) and non-defaulttemperature,top_portop_kvalues return errors. - Assistant message prefill returns an error.
- Computer use on the Claude API and Google Cloud needs the newer
computer_toolset_20260801tool. - Editing earlier turns invalidates thinking blocks, so conversations should be append-only.
- Responses can begin with a thinking block, so read content blocks by type, not by position. A safety classifier can also end a request with a
refusalstop reason.
What is still unknown. Anthropic gives no tokens-per-second figure for “fastest”. It does not publish the request mix behind the 75% average saving, or how a prompt of exactly 100,000 tokens is billed. The benchmark numbers are vendor-reported, and the Team credit’s rollover and allocation rules are not spelled out in the launch post.
Pre-launch claims did not hold up. Tracker sites projected $0.25 and $1.25 pricing and a 200K default context window, and social posts claimed Haiku 5.5 would beat Opus. The shipped model has a $0.10 and $0.50 price and a 1M window, and Anthropic’s own table shows Sonnet 5.5 ahead of it on every benchmark.
FAQ
When was Claude Haiku 5.5 released?
October 7, 2026. Anthropic first named it in the Opus 5.5 launch post on September 22.
How much does Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Cache reads start at $0.01 per million tokens, and the Batch API takes 50% off.
Why is it 75% cheaper on average if the rates are 90% lower?
The discount falls to 50% for prompts over 100K tokens, and the new tokenizer counts the same text as about 30% more tokens than Haiku 4.5 did.
Should I replace Sonnet 5.5 with Haiku 5.5?
Only for narrow, high-volume jobs whose output you can check cheaply. Sonnet 5.5 scores 70.6% against Haiku’s 39.2% on Terminal-Bench 4.0, and Anthropic still recommends Sonnet or Opus for complex agentic coding.
What is the model ID, and where can I use it?
claude-haiku-5-5 on the Claude API, with its own IDs on Amazon Bedrock (anthropic.claude-haiku-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Enjoy Worthview?
Add Worthview as a Preferred Source on Google to see more of our stories in Search.
Sethuram Kishore is the founder and editor of Worthview, an online publication established in 2008. With over 18 years of experience in SEO, digital marketing, and online publishing, he writes about AI, technology, business, and digital trends. He is also the founder of MoneyHulk, a personal finance and business publication.