Claude Haiku 5.5 vs Sonnet 5.5: Which Should You Use for Each Job?

Claude Haiku 5.5 vs Sonnet 5.5: Which Should You Use for Each Job?

Claude Haiku 5.5 launched on October 7, 2026 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, one tenth of Haiku 4.5’s price. The more useful question is not whether it is cheap, but which jobs it can now take away from Sonnet 5.5.

The short answer from the published numbers: Haiku 5.5 fits narrow, high-volume work such as classification, extraction, summarization, compaction and routing, plus lookups done by subagents under a larger model. Sonnet 5.5 keeps complex agentic coding and any task where a wrong answer is expensive. Anthropic says this itself, and its own benchmark table backs it up.

The price cut is real, but smaller than the headline. Anthropic puts the average saving at about 75%, not 90%, and the discount shrinks to 50% once a prompt passes 100,000 tokens. The rest of this piece shows where those limits sit and how to decide, job by job.

Haiku 5.5 vs Sonnet 5.5 at a glance

Sonnet 5.5 costs 20 times more per input or output token than Haiku 5.5 on prompts up to 100,000 tokens, and in return scores higher on every benchmark Anthropic published.

Haiku 5.5Sonnet 5.5
Input / output per 1M tokens$0.10 / $0.50 (up to 100K), $0.50 / $2.50 above$2 / $10
Cache read per 1M tokens$0.01 (up to 100K), $0.05 above$0.10
Latency tierFastestFast
Default effortmediumhigh
Context window / max output1M / 128K tokens1M / 128K tokens
Knowledge cutoffJune 2026June 2026
Anthropic’s positioningHigh-volume, latency-sensitive work such as classification, extraction and routingBest combination of speed and intelligence

What Anthropic announced

Anthropic calls Haiku 5.5 the cheapest, fastest and most capable small model it has released. It first named the model in the Opus 5.5 launch post on September 22 and shipped it on October 7, after Sonnet 5.5 on September 28. It is the first Haiku in about a year, and there was no Haiku 5.

SpecClaude Haiku 5.5
API model IDclaude-haiku-5-5
Context window1M tokens
Max output128K tokens (300K in the Batch API beta)
Input / outputText and images in, text out
ThinkingAdaptive, default effort medium
Knowledge cutoffJune 2026
PlatformsClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS
RetirementNot before October 7, 2027

It is also the first Haiku with an adjustable effort setting. Anthropic positions it for summaries, compactions, database queries and classification, as a subagent under Opus 5.5 or Sonnet 5.5 on coding work, and for speed-sensitive tasks such as live customer support and browser use. Its footnote adds that Haiku 5.5 is the fastest model at standard speed, but slower than Opus in Fast Mode.

Two other changes shipped the same day and matter for any cost comparison:

  • Sonnet 5.5 cache reads were cut by half, from $0.20 to $0.10 per million tokens. Anthropic says this makes Sonnet 5.5 about 20% cheaper on most agentic work.
  • Max and Team subscribers get a monthly API credit for the Claude Platform: $100 for Max 5x, $200 for Max 20x and up to $500 pooled for Team. The Python and TypeScript SDKs also gained beta support for computer use and browser use.

The price headline: 90%, 75% and the 100K cliff

Haiku 5.5 is priced in two tiers by prompt size, and the cheap tier is the one the headline describes.

Price per 1M tokensHaiku 5.5 (prompt up to 100K)Haiku 5.5 (prompt over 100K)Haiku 4.5
Input$0.10$0.50$1.00
Output$0.50$2.50$5.00
Cache read$0.01$0.05$0.10
Cache write (5 min)$0.125$0.625$1.25

The Batch API takes a further 50% off input and output. Thinking tokens bill as output, so a higher effort setting costs more per request.

Three details explain why Anthropic says “about 75% cheaper on average” instead of 90%:

  1. The 100K cliff. Above 100,000 tokens the rates are five times higher, so the saving versus Haiku 4.5 drops from 90% to 50%. Anthropic says about 90% of Haiku 4.5 requests fell below that line, but it does not publish how much of the spend did. A workload built on long documents or growing agent context will save less.
  2. The tokenizer. Haiku 5.5 uses the newer tokenizer shared with Sonnet 5.5 and Opus 5.5. The platform docs say the same text counts as roughly 30% more tokens than on Haiku 4.5, which Anthropic’s launch footnote describes only as “slightly more”. Your bill rises by that factor before the discount applies.
  3. Thinking is on by default. Adaptive thinking is enabled unless you turn it off, and thinking tokens bill as output, so the effort level feeds directly into cost.

To see what the tokenizer does, take a request with 2,000 input and 200 output tokens on Haiku 5.5. The same text would be about 1,540 and 155 tokens on Haiku 4.5. At published rates, one million such requests cost $300 on Haiku 5.5 and about $2,300 on Haiku 4.5, a saving near 87% rather than 90%. This is my arithmetic from the published rates and the 30% figure, and it ignores thinking tokens.

The practical takeaway: count your own tokens on the new tokenizer and check what share of your traffic exceeds 100K before you forecast savings.

One comparison point: OpenAI’s GPT-6 Luna lists the same $0.10 and $0.50 base rates, but its surcharge starts above 272,000 input tokens, not 100,000, according to VentureBeat’s check of official pricing.

What the benchmarks say, and what they hide

Haiku 5.5 beats Haiku 4.5 by a wide margin on every benchmark Anthropic published, and leads GPT-6 Luna wherever both appear. Sonnet 5.5 still wins every row. All figures below are vendor-reported, not independently verified.

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5Haiku as share of Sonnet
FrontierCode 1.1 (main)46.4%not reported42.4%52.1% (xhigh effort)89%
OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%86%
Humanity’s Last Exam (no tools)45.9%10.2%not reported56.9%81%
Chartography (no tools)46.4%6.4%29.1%61.6%75%
Terminal-Bench 4.039.2%0.0%16.4%70.6%56%
GDPval-AA v2.1 (Elo-style score)162073514371840not a ratio
AA-Briefcase v1.1 (Elo-style score)157861413361824not a ratio

The last column is my own calculation from Anthropic’s table. It shows where the gap to Sonnet is narrow and where it is wide. Haiku 5.5 gets within 11% to 14% of Sonnet on FrontierCode and OSWorld, but reaches only 56% of Sonnet’s Terminal-Bench score. That single gap is the clearest sign of where the cheaper model stops being a substitute.

Two caveats change how to read the table:

  • Effort setting. VentureBeat reports that Anthropic’s Terminal-Bench chart places the 39.2% score at maximum effort, with medium, the API default, near 20%. Check which effort level a benchmark used before comparing it with your own default-effort results.
  • Customer evidence is narrow. Box reports an 11-point gain over Haiku 4.5 at about half the latency, AlphaSense a 0.84 vs 0.76 score on 400 queries for a workload of about 8M calls a week, and HubSpot a 92.8% average on its CRM suite. Asana’s reported 30% latency drop and up to 2.5x faster inference per turn compare against an unnamed model. None of these give a tokens-per-second figure, and none compare across providers.

Anthropic’s own conclusion is consistent with the data: Sonnet 5.5 and Opus 5.5 stay the better choice for complex agentic coding, and Haiku 5.5 suits narrower tasks that were previously too expensive to run.

Which jobs go to Haiku, and which stay on Sonnet

Route by the job, not by the model’s reputation. The table starts from the use cases Anthropic names and the benchmark gaps above; rows marked “my judgment” are mine, not Anthropic’s.

JobRoute toWhy
Ticket tagging, routing, moderation queuesHaiku 5.5, low effortNarrow, high-volume work Anthropic names; classification and routing are the stated design target
Question answering over one or a few documentsHaiku 5.5, medium effortAlphaSense reports about 8M such calls a week and a 0.84 vs 0.76 score over Haiku 4.5
Summaries and context compaction inside agent loopsHaiku 5.5Named use case; cache reads cost $0.01 per million tokens at up to 100K
Subagent lookups under an Opus or Sonnet orchestratorHaiku 5.5Rogo’s example: a larger model builds the deck, Haiku pulls one revenue line from a 10-K
Browser and desktop automationHaiku 5.5, test first72.4% on OSWorld 2.1 vs 83.9% for Sonnet; SDK support is in beta
Multi-step coding and terminal agentsSonnet 5.5 or Opus 5.5Terminal-Bench 4.0: 39.2% vs 70.6%
High-stakes calls with no cheap check (my judgment)Sonnet 5.5 or aboveA wrong answer costs more than the token savings
Penetration testing and offensive security workNot Haiku by defaultIts safeguards block it; broader access goes through the Cyber Verification Program

The cost-per-success test

Per-token price is not the whole cost, because a cheaper model that fails more often costs more per correct answer. Take Terminal-Bench 4.0 as a stress case. Sonnet 5.5 costs $2 per million input tokens against Haiku 5.5’s $0.10 for prompts up to 100K, a 20x gap. Haiku solves 39.2% of tasks and Sonnet 70.6%.

If each attempt used the same number of tokens, which is a simplifying assumption, one Haiku success would cost about 2.6 Haiku-attempt units and one Sonnet success about 28. Haiku would be roughly 11 times cheaper per success, but only when you can detect a failure and retry. Above 100K tokens the price gap shrinks to 4x and the advantage drops to about 2x.

This points to the question that decides most routing: can you check the output cheaply? Extraction that must match a schema, classification you can sample against labels, and code that passes tests can all be verified, so Haiku plus a retry wins. A legal summary or a financial sign-off cannot be verified without a person, so the extra accuracy of a larger model is cheap insurance.

Four questions before you route a job

  1. Is the task narrow and well specified?
  2. Can a script, schema or sample check catch a wrong answer?
  3. Does the prompt stay under 100,000 tokens?
  4. Is a wrong answer cheap to fix?

Four yeses mean Haiku 5.5. One no means test it on your own data before moving. Two or more noes mean stay on Sonnet 5.5 or Opus 5.5. Opus 5.5 sits above Sonnet at $4 and $20 per million tokens, and Anthropic names both as better choices than Haiku for complex agentic coding, so the same four questions apply when choosing between Haiku and Opus.

Worked examples at published rates

These figures are my arithmetic from Anthropic’s list prices. They exclude thinking tokens, caching and the Batch API discount, and they assume both models see the same token counts.

ScenarioHaiku 5.5Sonnet 5.5Sonnet costs
A: 1M requests, 2,000 input and 200 output tokens each$300$6,00020x more
B: 100K requests, 150K input and 1K output tokens each$7,750$31,0004x more
C: same job as B, with prompts trimmed to 90K input tokens$950$19,00020x more

Scenario A is the clean case: the work stays under 100K tokens, so the full 20x price gap applies. Scenario B shows the cliff. The same model costs 8 times more than in scenario C because every token is billed at the higher rate once a prompt crosses 100K. Trimming prompts by 40% cut Haiku’s bill by 88%, but Sonnet’s by only 39%, because its listed price has no such threshold.

That makes prompt size a design choice worth testing: chunking a document or summarizing earlier turns can move a job under the line, provided accuracy holds. Anthropic’s launch materials do not say how a prompt of exactly 100,000 tokens is billed or which tokens count toward the threshold, so leave a margin and confirm with a test request.

The Sonnet price cut matters here too. Halving Sonnet 5.5’s cache-read price to $0.10 narrows the gap on cache-heavy agent loops. At up to 100K tokens, Sonnet’s cache read is now 10 times Haiku’s $0.01, down from 20 times.

Caveats before you switch

Safeguards are tighter for security work. Anthropic says Haiku 5.5’s cybersecurity safeguards are stricter than Haiku 4.5’s but looser than those on other recent models. They allow a wider range of defensive tasks than Sonnet 5.5’s safeguards, yet still block penetration testing and techniques more likely to be used by attackers. Its biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5. Organizations needing broader access can apply to the Cyber Verification Program or the Life Sciences Verification Program. On alignment, Anthropic reports far fewer instances of misaligned behavior and a lower willingness to cooperate with misuse than Haiku 4.5, with details in the system card.

Swapping the model ID is not enough. The docs list several changes from Haiku 4.5:

  • Manual thinking budgets (budget_tokens) and non-default temperature, top_p or top_k values return errors.
  • Assistant message prefill returns an error.
  • Computer use on the Claude API and Google Cloud needs the newer computer_toolset_20260801 tool.
  • Editing earlier turns invalidates thinking blocks, so conversations should be append-only.
  • Responses can begin with a thinking block, so read content blocks by type, not by position. A safety classifier can also end a request with a refusal stop reason.

What is still unknown. Anthropic gives no tokens-per-second figure for “fastest”. It does not publish the request mix behind the 75% average saving, or how a prompt of exactly 100,000 tokens is billed. The benchmark numbers are vendor-reported, and the Team credit’s rollover and allocation rules are not spelled out in the launch post.

Pre-launch claims did not hold up. Tracker sites projected $0.25 and $1.25 pricing and a 200K default context window, and social posts claimed Haiku 5.5 would beat Opus. The shipped model has a $0.10 and $0.50 price and a 1M window, and Anthropic’s own table shows Sonnet 5.5 ahead of it on every benchmark.

FAQ

When was Claude Haiku 5.5 released?

October 7, 2026. Anthropic first named it in the Opus 5.5 launch post on September 22.

How much does Haiku 5.5 cost?

$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Cache reads start at $0.01 per million tokens, and the Batch API takes 50% off.

Why is it 75% cheaper on average if the rates are 90% lower?

The discount falls to 50% for prompts over 100K tokens, and the new tokenizer counts the same text as about 30% more tokens than Haiku 4.5 did.

Should I replace Sonnet 5.5 with Haiku 5.5?

Only for narrow, high-volume jobs whose output you can check cheaply. Sonnet 5.5 scores 70.6% against Haiku’s 39.2% on Terminal-Bench 4.0, and Anthropic still recommends Sonnet or Opus for complex agentic coding.

What is the model ID, and where can I use it?

claude-haiku-5-5 on the Claude API, with its own IDs on Amazon Bedrock (anthropic.claude-haiku-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS.

Enjoy Worthview?

Add Worthview as a Preferred Source on Google to see more of our stories in Search.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.