§ News
By AI Blog Editor
Oct 8, 2026 · 14 min read
Claude Haiku 5.5 ships at a tenth of Haiku 4.5's price — then the new tokenizer eats a quarter of the cut back
On Oct 7 Anthropic shipped Haiku 5.5 at $0.10/$0.50 per million tokens up to 100k — ten times cheaper than Haiku 4.5 at the base, matching GPT-6 Luna. The new tokenizer eats about a quarter of the cut back.

On Wednesday October 7, 2026, Anthropic launched Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens on the first 100,000 tokens of each request, and $0.50 / $2.50 on everything above that threshold. The outgoing model, Haiku 4.5, cost $1 / $5 flat — so at the base tier the new price is ten times cheaper. The model is live on the Claude Platform, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Azure Foundry under the model ID claude-haiku-5-5. It is also the first Haiku-class model to ship with an adjustable reasoning dial — the Low / Med / High / Xhigh / Max levels familiar from the Opus and Sonnet lines.
Three things make this a Loop story rather than a routine model-card update. The headline price cut is real and matches GPT-6 Luna's base rate. The pricing page's own summary of that cut doesn't quite match the footnote arithmetic. And Simon Willison, who tested the preview inside twenty-four hours, found that the new tokenizer produces about 1.25 times as many tokens for the same prompt as Haiku 4.5 — a stealth price increase hiding inside the headline price cut.
The price cut, measured three different ways
Anthropic's launch page describes Haiku 5.5 as costing "around 75% less to run" than Haiku 4.5. The same page's footnote splits that into a 90% reduction up to 100,000 tokens per request and a 50% reduction beyond. Those two figures don't average to 75% in any natural way — the real answer depends entirely on how your workload is distributed around the 100k threshold.
Take the numbers at face value. For input tokens Haiku 4.5 charged $1 per million at every token count. Haiku 5.5 charges $0.10 up to 100k and $0.50 beyond. A workload that always stays under 100k gets a 90% cut. A workload that averages 500k per request gets closer to the 50% number. The "around 75% less" headline is defensible only for a specific middle regime, and the page doesn't say which. For a buyer trying to model the migration, it's a figure in search of a workload.
Then there is the tokenizer. Willison ran the same long prompt against both models through his Claude token counter and found Haiku 5.5 produced about 1.25 times as many tokens as Haiku 4.5. New tokenizer, new vocabulary, same text. On a base-tier workload where the sticker price fell 90%, a 25% token inflation lands the effective cut closer to 88%. On the above-100k tier where the sticker price fell 50%, the same inflation brings the effective cut down to roughly 38%. Still a meaningful reduction. Not the one in the headline.
Anthropic did not disclose the tokenizer change in the launch post. It is also not surfaced on the pricing page. A customer migrating a production workload from Haiku 4.5 to 5.5 today, estimating their new bill by multiplying their current token volumes by the new per-million rate, would be systematically off by about 25% in Anthropic's favour. That is a sentence someone's finance team should read before the next quarterly review.
The GPT-6 Luna match is narrower than it looks
The $0.10 / $0.50 base rate matches GPT-6 Luna exactly. On every launch benchmark Anthropic published — GDPval-AA v2.1, OSWorld 2.1, Humanity's Last Exam, Terminal-Bench 4.0, FrontierCode, Chartography — Haiku 5.5 posts higher scores than Luna. If you can hold your workload under 100,000 tokens per request, the choice between the two is a benchmark story rather than a pricing story.
Push past 100,000 tokens and the match collapses. OpenAI's Luna raises its price only above 272,000 tokens, and the above-threshold rate is $0.20 input / $0.75 output — less than half what Haiku 5.5 charges above 100k. For long-context workloads — document analysis, codebase ingestion, long conversation summarisation — Luna is materially cheaper. For short-form inference — classification, extraction, routing, lightweight agent steps — Haiku 5.5 is competitive or ahead. The pitch deck line is "we match Luna," which is true at the base. The engineering truth is "we match Luna for a narrower band of workloads than the headline suggests."
The reasoning dial complicates the comparison further. Haiku 5.5 is the first small-and-cheap Anthropic model with the Low/Med/High/Xhigh/Max selector that was previously an Opus and Sonnet feature. Reasoning tokens are output tokens, so turning the dial up multiplies your bill. Willison's armadillo-in-fishnet-tights SVG test — yes, this is now a standard industry benchmark — ran about 0.09 cents at Low effort and 3.4 cents at Max, a nearly 40x spread on a single pelican-adjacent prompt. The reasoning dial is powerful. It is also an invitation to spend forty times what you thought.
The subscriber-API bundle is new
Buried in the same launch post is a structural change to how Anthropic sells its API: "Second, this week, we'll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform." Max 5x subscribers get $100 per month of API credit. Max 20x subscribers get $200. Team seats pool up to $500 per month across the team's users. Unused credit does not roll over to the next month. Auto-recharge can be turned off, which means the requests stop when the monthly budget is spent rather than quietly billing a credit card.
This is the first time Anthropic has bundled API consumption into the subscription product. The economic effect is that Max 20x — $200 a month retail — becomes a $200-of-API tier with Claude Code and the chat product thrown in for free, or a $200-of-chat tier with $200 of API thrown in for free, depending on how a given user actually consumes it. Either framing is bad for the per-usage API competitor whose entire pitch is pay-only-for-what-you-use. OpenAI's ChatGPT Plus and Enterprise plans still sell inference as a line item billed against the organisation's API key separately from the subscription. Google's AI Pro and AI Ultra plans bundle Gemini Pro inside the chat product but don't extend credit to the paid API.
The Sonnet 5.5 cache-read reduction, also announced in the same post, cuts the per-million-token cache-hit price from $0.20 to $0.10. For long-running agents and iterative RAG workflows that depend on cache hits, this is the second substantial pricing concession of the week. Combined with the Haiku 5.5 base-tier cut and the subscription credit, the Oct 7 post is Anthropic's largest single-day price action since the Opus 5.5 launch on Sept 22, when Opus 5.5 dropped to $4 / $20 — 20% below Opus 5's rates.
Why three cuts in sixteen days
The pattern is now a sequence. Sept 22: Opus 5.5 at a 20% cut. Sept 28: Sonnet 5.5 with a tokenizer change and price hold. Oct 7: Haiku 5.5 at a 10x base-tier cut, Sonnet 5.5 cache-read halved, API credit bundled into subscriptions. Three pricing actions in sixteen days, each one framed as a feature of the model release rather than as a response to pressure.
The pressure is reasonable to infer. The Ramp AI Index numbers that leaked via The Decoder in early October had Anthropic at 51% of enterprise AI token spend against OpenAI's 44.5% — a sharp jump from May's 34.4%. The way to defend 51% is to make the per-token cost of staying on Anthropic indistinguishable from the per-token cost of leaving. Matching GPT-6 Luna at the base, cutting Sonnet 5.5 cache reads to a dime, and giving Max subscribers $100-$500 of monthly API credit are three independent ways to do that.
The way to defend against defection to the open-weight Chinese labs — DeepSeek, Qwen, GLM, and now Mistral Large 4 at the end of this month — is harder. Price-cutting a frontier model to zero is still more expensive than running an open-weight model you host yourself. Anthropic's answer there is a product story (reasoning dials, 1 million token context, cached workloads) rather than a pricing story. On the hosted-API axis, the Oct 7 bundle is Anthropic saying we will not lose this on price alone. Whether it works turns on whether the subscriber-API bundle is read by enterprise buyers as a discount or as a lock-in.
What to watch
- The Haiku 5.5 context window. Anthropic's launch page does not state it. The pricing tiers are structured around 100,000 tokens, which strongly implies the window is at or above that figure — probably 200,000 like Haiku 4.5. If it turns out to be smaller, the "matches Luna" claim weakens further.
- Independent reproduction of the Oct 7 tokenizer finding. Willison's 1.25x number is one data point from one tester. If three or four large customers (Vercel, Cursor, Zed, Codeium — all of whom have published Claude pricing analyses before) report similar inflation ratios, the "around 75% less" headline becomes a problem Anthropic needs to clarify. If the ratio turns out to be workload-dependent in ways nobody has mapped, the model's effective price becomes "whatever your prompts happen to look like."
- Whether OpenAI responds on Luna's above-272k pricing. The current gap at long context is OpenAI's structural advantage. If Luna drops below $0.20 / $0.75 at the high tier in the next thirty days, the Haiku 5.5 base match is undone. If it holds, Anthropic has staked its small-model business on short-form workloads and will defend that ground.
- Whether the API-credit bundle produces measurable Claude-Platform-vs-ChatGPT-Plus churn. Max 5x and Max 20x subscribers now have an incentive to migrate agent and API workloads onto the Claude Platform rather than onto a separately-billed OpenAI or Google key. Ramp AI Index numbers for November will show whether that incentive moves spend. If the Nov report has Anthropic above 55%, the Oct 7 bundle worked. If it holds at 51%, the bundle neutralised a defection that was about to happen rather than causing a gain.
Three pricing actions in sixteen days from a lab that still carries a $200 billion valuation. The headline is "ten times cheaper." The arithmetic is "about four times cheaper after the tokenizer." Both are true. The gap between them is where the next finance-team conversation at a mid-sized Anthropic customer happens.
* * *
Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.
Elsewhere in this issue
3 moreLetters
Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.