# Claude Haiku 5.5 Is Up to 90% Cheaper. Here Is the Catch.

Published 11 October 2026 by Kaizen AI Team. Tags: AI Tools, Small Business.

Page: https://kaizenaiconsulting.com/claude-haiku-5-5-cheaper-catch/

Claude Haiku 5.5, Anthropic's cheapest model, launched on 7 October 2026 at up to 90% less than the model it replaces. We have been through the pricing, and our view is simple: it is a genuine price cut, but the headline number is not the number you will see on your bill.

## What actually changed

Haiku is the small, fast tier of Claude, the one you point at repetitive work: sorting emails, tagging enquiries, drafting first-pass replies, pulling details out of invoices. Until last week, Claude Haiku 4.5 cost $1 per million input tokens and $5 per million output tokens.

According to [VentureBeat's coverage of the launch](https://venturebeat.com/technology/anthropic-launches-claude-haiku-5-5-with-90-api-price-reduction-matching-gpt-6-luna), Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is a 90% cut on the shorter tier. It also matches the price of OpenAI's GPT-6 Luna, which launched on 22 September, so the two vendors now sit at exactly the same cheap-tier price.

If tokens are a new word to you: they are the chunks of text an AI model reads and writes, roughly three quarters of a word each. You pay per token, so cheaper tokens mean cheaper automations.

## Catch one: the 90% only applies to shorter prompts

Above 100,000 tokens in a single request, the price steps up to $0.50 input and $2.50 output per million. That is still half the old price, but it is five times the short-prompt rate. [The Decoder](https://the-decoder.com/claude-haiku-5-5-arrives-with-massive-price-cuts-proving-the-ai-pricing-arms-race-is-far-from-over/) reports that Anthropic says the short tier covers about 90% of Haiku 4.5 requests, so most everyday jobs will land in the cheap band. But if you feed a model a whole year of paperwork in one go, you will cross the line without noticing.

## Catch two: the new tokenizer counts more tokens

The tokenizer is the part of the model that chops your text into tokens. Haiku 5.5 uses the newer one, and press coverage puts it at roughly 30% more tokens for the same text compared with Haiku 4.5. The same email now costs more tokens, so the real saving is smaller than 90%. Anthropic's own headline, as reported by the same outlets, is closer to 75% cheaper on a typical mix of requests once this is factored in.

We think that is the honest number to plan around, and even that depends on your workload.

## A worked example, with the assumptions on show

Take a small business sorting 1,000 customer emails a month. Assume each email plus instructions is 500 tokens going in, and the model writes a 100 token reply or tag coming out. These figures are illustrative, not a measurement of any real business.

- **On Haiku 4.5:** 0.5 million input tokens at $1 is $0.50, and 0.1 million output tokens at $5 is $0.50. Total: about $1.00.
- **On Haiku 5.5:** allow 30% more tokens, so 0.65 million in at $0.10 is $0.065, and 0.13 million out at $0.50 is $0.065. Total: about $0.13.

That is roughly 87% lower, which is close to the headline, because this job sits well under the 100,000 token line. A job that regularly sends long documents would save less.

The bigger point is that both numbers are tiny. For a thousand emails a month, the AI bill was never the issue. The price cut matters when you run something at volume: every enquiry, every invoice, every review, all day.

## Cheaper is not the same as better for the job

A small model is a small model. It is built for routine, well-defined work, and the cheapest tier can need more retries or more checking than a mid-tier model on anything that needs judgement. If a cheap model gets a customer reply wrong and a person has to fix it, the saving has gone. We covered the idea of matching the model to the job in our Playbook posts: thinking work goes to a top-tier model, routine work goes to the cheap one, and a person checks anything a customer will see.

It is also worth knowing that Haiku 5.5 is the first Haiku with an adjustable effort setting, according to [MarkTechPost](https://www.marktechpost.com/2026/10/07/anthropic-releases-claude-haiku-5-5-a-small-model-with-1m-context-priced-at-0-10-per-million-input-tokens/). Turning effort up makes it think harder on a tricky request, at a higher token cost. Cheap is a dial, not a fixed point.

## What we would do with this

This is what we would tell a client who asked us on Monday morning whether to change anything:

1. **Do nothing if you only use the chat apps.** A subscription price does not move because an API price did. This change is for anything connected through the API, such as an automation, a chatbot or a tool built on top of Claude.
1. **If you pay for an automation, ask what model it runs on.** If it is Haiku 4.5, ask your supplier whether it has moved to 5.5 and whether the saving has been passed on. Do not assume it has.
1. **Pick one routine job and test it.** Run the same 50 real examples through both models, and compare the cost and the quality side by side. Cost per job is the only figure that matters, not cost per token.
1. **Keep a person on anything customer-facing.** A cheaper draft that nobody reads is more expensive than it looks.
1. **Watch the 100,000 token line.** If your prompts are long, work out which price band you are really in.

## The bigger picture

In under three weeks, the cheap tiers of two of the biggest AI vendors have landed at the same price: $0.10 in, $0.50 out per million tokens. That is a sign of how fast the floor is dropping. What it means for a small business is that the cost of running routine AI work is becoming a rounding error, and the hard part is no longer paying for the model. The hard part is deciding which job to give it, writing down what good looks like, and checking the output.

That is where most firms stall. The ONS found only about one in ten UK firms using AI use it extensively, and the tokens were never what held them back. Cheaper tokens will not fix a missing process.

The price per token went down 90%. The price of a badly chosen job did not move at all. If you want a second pair of eyes on which of your routine jobs is worth automating, and which model suits it, that is exactly what our free 30 minute AI audit is for.
