AI & Machine Learning

Anthropic Launches Claude Haiku 5.5 With Prices Up to 20x Below Sonnet

Claude Haiku 5.5 cuts costs for everyday AI tasks, with API rates up to 20x below Sonnet and Anthropic-reported benchmark gains over GPT-6 Luna.

By Daniel Mercer Edited by Samantha Reed Published: Updated:
Anthropic Launches Claude Haiku 5.5 With Prices Up to 20x Below Sonnet
Anthropic launches Claude Haiku 5.5 for lower-cost everyday tasks, with standard token rates up to 20x below Sonnet. Photo: Planet Volumes / Unsplash

Key Notes

  • Claude Haiku 5.5 targets summaries, document questions and supporting agent tasks.
  • Standard input and output rates are 20x below Sonnet for prompts up to 100,000 tokens.
  • Anthropic reports benchmark gains over GPT-6 Luna, while more complex work still favors larger models.

Anthropic has released Claude Haiku 5.5, a new small model aimed at making everyday AI work faster and less expensive. The October 7 launch targets summaries, document questions, database queries and other recurring tasks, including supporting roles inside agents powered by larger Claude models.

The headline price difference is substantial: for prompts up to 100,000 tokens, standard input and output rates are 20x lower than Sonnet 5.5’s. That comparison concerns API token prices, rather than subscription fees or a guarantee that every completed task will cost 20x less.

How Much Does Claude Haiku 5.5 Cost?

Anthropic’s pricing table separates shorter and longer prompts. Once a prompt exceeds 100,000 tokens, Haiku 5.5’s input and output rates rise fivefold, reducing the price advantage over Sonnet. Applications that regularly send extensive documents should account for that threshold before estimating savings.

Token type Haiku 5.5: up to 100k Haiku 5.5: over 100k Sonnet 5.5
Input $0.10 $0.50 $2.00
Output $0.50 $2.50 $10.00
Cache reads $0.01 $0.05 $0.10

All figures are per million tokens. At the shorter-prompt rates, a workload consuming 10 million input tokens and 2 million output tokens would cost $2 on Haiku, compared with $40 on Sonnet, before caching, additional services or differences in token consumption. This is a calculation using equal token volumes, not a measured comparison of equivalent work.

Moving from Haiku 4.5 also requires more than substituting prices in a spreadsheet. Anthropic’s migration guide says the newer tokenizer produces approximately 30% more tokens for the same text, with variation by content. Developers should recount representative inputs and revisit output limits; a cheaper token does not necessarily represent the same amount of text.

Where Haiku Fits Alongside Opus and Sonnet

Haiku’s role is easiest to understand as a division of labor. A larger model can manage an involved project while a smaller model handles a bounded assignment, such as extracting a figure from a document or summarizing material for the next step. The useful question is whether the cheaper model completes that assignment reliably enough to avoid expensive retries and manual correction.

Consider an internal document assistant. Retrieving a short passage and identifying the requested number is a different workload from reconciling conflicting evidence across a large archive. Testing those separately would give a team a clearer basis for routing requests than treating every document question as equally difficult.

Our coverage of subscription value examined a different economic question: how much usage a monthly plan provides. Haiku’s API pricing matters to applications billed by consumption. A subscription comparison cannot establish the operating cost of an application making thousands of model calls.

What the GPT-6 Luna Comparison Shows

In Anthropic’s published evaluations, Haiku 5.5 scores 72.4% on the OSWorld 2.1 offline subset, against GPT-6 Luna’s 48.9%. On Terminal-Bench 4.0, the reported figures are 39.2% and 16.4%. Those are vendor-reported results under the listed test conditions, rather than independent confirmation that Haiku wins on every workload.

The same results also show why the cheaper model does not replace the rest of Claude’s range. Sonnet 5.5 reaches 70.6% on Terminal-Bench 4.0 in that comparison. For a complex coding task, paying more per token can still be worthwhile if the stronger model finishes successfully with fewer attempts.

What Developers Should Check Before Switching

The migration documentation identifies adaptive thinking and response handling as important integration changes. Applications should select response blocks by type, rather than assume the first block contains the final answer. Existing sampling settings and assistant-prefill patterns also need review before an older integration is moved over.

A practical rollout would compare accuracy, completion time, retries and total spending on a fixed set of real requests. Include both typical short prompts and examples that cross the pricing threshold. That makes it possible to judge the cost of a successful result, rather than optimize a token rate while overlooking failures elsewhere in the workflow.

Haiku 5.5 is available through Anthropic and its major cloud platforms. The immediate opportunity is to make narrowly scoped AI work economical at larger volumes. Whether it belongs in a particular application depends on how well it handles that application’s tasks, with the larger Claude models remaining options when the work demands more capability.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Enterprise Tech, News