bytevyte
bytevyte
Language

Anthropic's Claude Haiku 5.5 Pricing Cut Reaches 90% as Labs Fight on Cost per Task

The Claude Haiku 5.5 pricing cut reaches 90% below 100K tokens, matches GPT-6 Luna and shifts the labs' fight to cost per task.

Claude Haiku 5.5 pricing
Photo by Brecht Corbeel on Unsplash

Anthropic has repriced its cheapest production model to a level that lands directly on top of OpenAI's budget tier. Claude Haiku 5.5 now costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, a 90% reduction against the $1/$5 rates of Haiku 4.5. The Claude Haiku 5.5 pricing change, published with the model's launch this week, applies to what Anthropic says is roughly 90% of the requests previously routed to the older Haiku.

Prompts above the 100,000-token mark are billed at $0.50 and $2.50 per million. Anthropic puts the blended saving on real workloads at about 75%, lower than the headline figure because it folds in request mix and a new tokenizer that consumes slightly more tokens per task. The model is available on Amazon Bedrock, Claude Platform on AWS and AWS GovCloud (US) for regulated workloads.

The 100,000-Token Threshold Is the Actual Product Decision

A single per-token price would have been simpler. Anthropic chose a two-tier structure instead, and the boundary sits where agentic workloads tend to accumulate.

A coding agent that carries tool output across a long session crosses 100,000 tokens without difficulty, and past that point input costs multiply fivefold. The discount is real for the summarisation, classification and tool-query jobs that dominate high-volume traffic. It narrows for the long-context jobs that agent frameworks generate once a session runs long enough.

TierInput ($/M tokens)Output ($/M tokens)
Claude Haiku 4.51.005.00
Claude Haiku 5.5, prompts up to 100K0.100.50
Claude Haiku 5.5, prompts above 100K0.502.50

The first-ever effort control on a Haiku-class model changes the arithmetic further. Teams can set reasoning depth per request, with medium as the default, trading latency and token consumption against answer quality. That moves cost control from model selection down to the individual call, which matters for pipelines that mix trivial and demanding steps in the same run.

The 1M-token context window is the other half of the picture. It expands what a single Haiku call can hold, and it makes the 100,000-token boundary easier to cross. Teams that adopt the larger window for retrieval-heavy work will pay the higher tier rate on the very tokens the window encourages them to send. The window is useful for long-document analysis, but the tiered rate means filling it is not cheap.

Sonnet 5.5 gets a matching adjustment. Anthropic halved its cache-read price to $0.10 per million tokens, which it calculates makes Sonnet 5.5 roughly 20% cheaper on most agentic work. Cache reads are the repeated context that agents resend on every turn, so the cut lands on the same workload category as the Haiku repricing.

Cache pricing behaves differently from standard tokens because cached context is written once and read many times. Halving the read rate compounds across every step of an agent loop. For a pipeline that replays a 50,000-token system prompt across a hundred turns, the Sonnet 5.5 cut reduces the dominant line item rather than a marginal one, which is why the effective saving on agentic work runs ahead of the raw change to the cache-read rate.

Why Claude Haiku 5.5 Pricing Moved the Fight Down-Stack

Haiku 5.5 is the third Claude 5.5-family model in about a month, with Fable 5.5 still outstanding. The cadence matters less than the direction of the cuts. The 90% reduction lands on the cheapest tier, not the flagship.

That is where enterprise AI budgets actually scale. A frontier model is evaluated once a quarter and used on a bounded set of hard problems. A small model runs inside every subagent loop, every retrieval step and every classification call, and its cost grows linearly with usage. Cutting the flagship price by 10% saves a rounding error against cutting the workhorse tier by 90%.

Matching OpenAI's GPT-6 Luna pricing on the entry tier is the clearest signal of intent. The two labs have converged on the same number for the same workload, which suggests the entry-level rate is set by competitive response rather than by cost-plus calculation. Neither vendor has an obvious margin cushion left at $0.10 per million input tokens, so the next round of competition is more likely to arrive as capability per token than as a lower headline rate.

Benchmarks, and What They Do Not Settle

Anthropic positions Haiku 5.5 as its most capable small model, and the published numbers support a step up from Haiku 4.5 across coding, tool use and computer use. On OSWorld 2.1 computer-use tasks it scores 72.4%, against 48.9% for GPT-6 Luna. On Humanity's Last Exam it reaches 45.9%. Anthropic frames the model as a sharp jump over Haiku 4.5 on coding, tool use and computer use alike.

Capability parity at the entry tier is a bigger deal than the price parity. If a small model can drive a browser or complete a multi-step tool call at 72.4%, the case for routing that work to a mid-tier model weakens, and the token bill falls twice: once from the rate cut, once from the model substitution.

The 45.9% on Humanity's Last Exam is the more revealing figure, because that benchmark targets expert-level reasoning rather than routine automation. A small model clearing roughly half of it is competitive on work that was reserved for larger models a year ago, and that is the substitution pressure mid-tier offerings now face.

The open question is measurement. Effort controls mean the same model produces different cost and quality profiles depending on a setting most teams will not tune per endpoint. Benchmark scores published at default medium effort will not predict the cost of a production pipeline that runs at low effort for volume and high effort for exceptions.

What to Verify Before Migrating

The 75% average saving is an Anthropic estimate built on a workload mix, not a guarantee for any single application. Teams should model their own request distribution against the 100,000-token boundary before assuming the headline rate applies. A service that sits mostly under the threshold will capture close to the full 90%. One that runs long agent sessions will capture considerably less.

Two-tier billing also complicates forecasting. Cost per request is no longer a fixed multiple of token count, and the boundary creates an incentive to split work across calls to stay under the threshold, which adds orchestration overhead and can reduce answer quality. Whether that trade pays off depends on how tightly a workload hugs the line.

The new tokenizer is the second variable. It consumes slightly more tokens for the same input than the tokenizer behind Haiku 4.5, so a per-token price cut does not translate one-for-one into a per-task cost cut. Anthropic's own 75% figure already accounts for this; a migration plan built on 90% does not.

Anthropic also introduced a monthly API credit for Claude Max and Team subscribers, which lowers the effective cost for teams already on those plans. It names customers including Asana, HubSpot, AlphaSense, Box and Cognition as users of the release, a list skewed toward products where per-request cost determines whether a feature ships at all. The Claude Haiku 5.5 pricing structure is therefore worth testing against a real workload rather than a benchmark score.

Why This Matters

The Claude Haiku 5.5 pricing move confirms that the frontier labs' competitive centre has shifted from benchmark supremacy to cost per completed task, because that is the unit enterprise budgets are built on. For buyers, model selection is now a cost-engineering exercise: the same workload can run at materially different prices depending on where it sits against the 100,000-token line and what effort setting it uses. The next thing to watch is whether Fable 5.5 arrives with the same treatment, since a price cut at the top of the family would signal that the labs intend to compress margins across the stack rather than only at the entry point.

Sources

Introducing Claude Haiku 5.5

Introducing Claude Haiku 5.5 on AWS

Claude Haiku \ Anthropic

Photo by Brecht Corbeel on Unsplash

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.