Claude Haiku 5.5 launches as Anthropic's fastest and cheapest small model

Claude Haiku 5.5 is Anthropic's new small model for high-volume and latency-sensitive work, with the company claiming roughly 75% lower average run cost than Haiku 4.5.

Mason Reed

Anthropic has launched Claude Haiku 5.5, describing it as the fastest, cheapest and most capable small model the company has released so far.

The October 7 release is aimed at high-volume and latency-sensitive work rather than replacing the larger Opus and Sonnet models outright. Anthropic says Haiku 5.5 can handle quick repetitive tasks such as summaries, classification and database queries, while also working as a lower-cost coding subagent beside Opus 5.5 or Sonnet 5.5.

Anthropic says average run cost drops by about 75%

The headline efficiency claim is substantial but needs the vendor label attached. Anthropic says Haiku 5.5 costs around 75% less on average to run than Haiku 4.5.

Its published pricing is tiered by context length. For prompts up to 100,000 input tokens, Anthropic lists $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 input tokens, the rates rise to $0.50 and $2.50 respectively.

Anthropic says that works out to a 90% price reduction for shorter-context use compared with Haiku 4.5 and a 50% reduction for longer prompts. The roughly 75% figure is the company's estimate of average run cost, not an independent workload study.

The company also describes Haiku 5.5 as its fastest model. Anthropic's own footnote makes a useful qualification: that comparison excludes Opus models running in Fast Mode.

Haiku 5.5 is available on Anthropic's own platform as well as Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic also says this is the first Haiku model with an adjustable effort setting, giving developers another lever between cost and capability.

The launch also cuts some Sonnet costs

Haiku is not the only pricing change. Anthropic is halving Claude Sonnet 5.5 cache-read pricing, from $0.20 to $0.10 per million tokens. It says that can make most agentic Sonnet 5.5 work around 20% cheaper, again as a company estimate tied to its assumed usage patterns.

Anthropic is also adding monthly Claude API credit for some subscribers: $100 for Max 5x, $200 for Max 20x and up to $500 pooled for Team plans. The credit is intended for building agents and applications on the Claude Platform rather than being a general cash rebate.

The family positioning matters because Anthropic is not claiming Haiku 5.5 should replace Opus or Sonnet for every hard task. Its pitch is that a small, fast model can take more of the routine volume, including subagent work inside larger coding workflows, while the larger models handle tasks where extra capability justifies the higher cost.

Anthropic's newsroom provides a second first-party check on the October 7 launch and describes Haiku 5.5 in the same high-volume, cost-sensitive terms. That corroborates the release state without turning Anthropic's benchmark or cost claims into independent verification.

There will still be a gap between vendor benchmarks and real production economics. Prompt length, cache reuse, tool calls and failure/retry rates can matter as much as the sticker rate on a model card.

What changed this week is simpler: Anthropic has pushed the Haiku tier down in price, tied it more explicitly to high-volume agent work, and cut a common cached-input cost for Sonnet at the same time.

Sources

More from Xarmo News