§ News
By AI Blog Editor
Jul 2, 2026 · 17 min read
The price sheet didn't move — Anthropic launched Claude Sonnet 5 with Sonnet 4.6's rate card and a new tokenizer that makes the same English text cost forty percent more
On June 30, 2026 Anthropic previewed Claude Sonnet 5 at the same $3/$15 per-million-token rate as Sonnet 4.6. The tokenizer underneath the rate card is new, and independent testing on the launch day showed it burns roughly 40% more tokens on English text.

On Tuesday June 30, 2026, Anthropic previewed Claude Sonnet 5 with a pricing table identical to the one Sonnet 4.6 shipped under: $3 per million input tokens, $15 per million output tokens. There is even an introductory promo through August 31 — $2 / $10 — that the launch post frames as a discount on top of a discount. The blog copy calls the model "our most efficient way to run agents at scale."
Everything on that page is technically true. What the page doesn't dwell on, and what Simon Willison pulled apart the same afternoon, is that the tokenizer underneath the rate card is new. The nominal price per token is unchanged. The number of tokens the model reads and writes is not.
What actually costs what
Willison ran the Universal Declaration of Human Rights through both models. Sonnet 4.6 encoded it in 2,356 tokens. Sonnet 5 encoded the same text in 3,341 tokens. That is a 1.42× increase on identical input. He then ran a 4,279-line Python file through the same test: 44,014 tokens on Sonnet 4.6, 56,113 tokens on Sonnet 5 — 1.27×. Spanish documents came out at 1.33×. Only Simplified Mandarin, at 1.01×, was materially unchanged.
Anthropic's own launch page, in a footnote, admits the range: the tokenizer "changes how the model processes text… the same input can map to roughly 1.0–1.35× depending on the content type." The 1.35× ceiling in the launch note is smaller than the 1.42× Willison found on English prose — either because Anthropic's internal tests weighted code more heavily than long-form text, or because the marketing team rounded down. Either way, the two numbers describe the same reality: on the substrate most enterprise agents actually run over, Sonnet 5 costs about thirty to forty per cent more per query than Sonnet 4.6 at the same nominal rate.
That is a price hike. It is not on any of the launch slides. If a SaaS vendor did this — held the sticker price steady and quietly re-metered the counter — the trade press would use the phrase dark pattern. Because it is a tokenizer swap by a frontier lab, the trade press used the phrase cheaper way to run agents.
The benchmarks are real; the pricing spin is not
The capability numbers behind the launch are, for what it's worth, real. Sonnet 5 scores 92.4% on SWE-bench Verified, twelve points above Opus 4.6. On OSWorld-Verified, the computer-use evaluation, it clears 88.3% — sixteen points above the 72.4% human-expert baseline the benchmark was calibrated against. GPQA Diamond lands at 96.2%. ARC-AGI-2 at 84.7%. The context window is 1M tokens, with a 128K output cap, and adaptive thinking is on by default unless the caller disables it.
Those are Opus-tier numbers on a model that costs one-sixth of Opus 4.8's headline rate. That is a real product move. The story is not that the capability isn't there. The story is that Anthropic could have priced the capability jump into the rate card, and instead priced it into the tokenizer. One is legible. The other requires a Simon Willison and a spreadsheet to explain to your finance team.

The competitive box the tokenizer solved
Sonnet 5 landed four days after OpenAI's June 26 GPT-5.6 preview — Sol at $5 input / $30 output, Terra at $2.50 / $15, Luna at $1 / $6, all covered in the Loop's Sol writeup. Terra hits the same headline rate as Sonnet 5's introductory tier. Gemini 3.5 Flash, per TechCrunch, sits below all of them on the per-token sticker. Anthropic walked into a rate-card war it could not win at the top of the page.
So it didn't try. The rate sheet is unchanged, the marketing angle is "cheaper way to run agents," and the actual cost per unit-of-work moves through the tokenizer instead. That is a smarter play than dropping the headline number by 10% — enterprise procurement compares dollars-per-million on a spreadsheet, not tokens-per-document on a graph — but it only holds until customers wire up their own accounting.
Which is what Willison did on day one. And what every finance-ops team running an agent stack is going to do by the end of the week. The half-life of the effective-price story is until the first monthly invoice.
The half-truth in the "cheaper agents" pitch
The Engadget framing is the version worth taking seriously. Sonnet 5, their coverage argues, addresses the specific enterprise pain point that "agentic AI tools generate vastly more queries than humans" and are "largely responsible for instances where customers, particularly enterprise businesses, blow huge amounts of their budget on tokens." The efficiency story assumes the workload shape is agents, not documents. And for pure agent traffic — short structured prompts, tool calls, JSON responses — the tokenizer overhead is closer to the 1.0× floor than the 1.42× ceiling.
That is the version of the pitch that survives contact with a real customer. If your load is coding agents making short API-shaped calls, Sonnet 5 does deliver Opus-tier capability at a fraction of Opus pricing, and the tokenizer overhead is close to noise. If your load is document-heavy — RAG over English text, long-context summarisation, retrieval on multilingual corpora — you just took a forty-per-cent rate increase and got told it was a discount.
The customer who reads the launch post and picks the model without measuring their own workload is the customer this pricing was designed for. That is a sentence that costs $500.
The Opus, Mythos, and Sonnet 4.6 question
There is a footnote worth pulling out. Anthropic tells enterprise customers that the safety envelope for Sonnet 5 is "significantly less capable at cyber tasks than Mythos 5" — Mythos 5 being the model that got export-controlled on June 12 and cleared to about a hundred US institutions on June 26. That is a policy sentence, not a capability boast. It is the sentence a Commerce Department lawyer wants to see in the launch post so the model doesn't attract a fresh directive on day two.
The pricing implication is quieter and more consequential. If Sonnet 5 does what Opus 4.8 did four months ago at one-sixth of Opus 4.8's rate, Opus becomes the model whose only remaining edge is the small subset of frontier-only workloads customers now think twice about routing through it. That is Anthropic cannibalising its own top of stack — on purpose, because the alternative is watching Gemini 3.5 Flash and GPT-5.6 Terra eat the middle. Sonnet 4.6 is the model getting deprecated in the same breath. That is the sentence customers on annual contracts are about to read.
What this means
Tokenizer swaps are the new price move. Anthropic held the rate card, moved the counter, and shipped the composite as a discount. That template is now on the shelf. Expect OpenAI and Google to reach for it on the next mid-cycle refresh, and expect FinOps dashboards to grow a "tokens-per-document" column by the end of Q3.
The "cheaper agents" framing is workload-dependent. For short-call agent traffic the pricing story is real. For document-heavy workloads it inverts. The launch post does not distinguish between the two. Every Sonnet 5 rollout should now start with a bake-off on your own corpus, not on Anthropic's benchmark suite.
Opus 4.8's role just narrowed. Sonnet 5 at Sonnet 4.6 pricing hitting Opus-tier benchmarks compresses the top of the Anthropic stack to whatever the Preparedness Framework flagship needs to be for the next Commerce directive. Anthropic is not selling Opus as the general-purpose flagship any more. It is selling it as the model you route to when Sonnet 5 refuses.
The tokenizer footnote is the story the launch post did not want. Willison, an afternoon, and a single Python file were enough to turn the "unchanged pricing" claim into a forty-per-cent increase for the workloads customers actually run. The next launch, whoever ships it, will get audited within twenty-four hours. Vendors that don't publish the tokenizer delta up front will be assumed to have hidden it.
Anthropic priced this model for the customer who reads the top of the launch post. The customers who read the footnotes are about to hit send on procurement Slack messages that begin hey, so about the token bill. Both of those customers are still Anthropic's. The difference is whether the discount is a discount, or the sticker held because the meter changed underneath it.
* * *
Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.
Elsewhere in this issue
3 more- 01
News
The means of production — Palantir posted $1.94 billion in one quarter, then used the shareholder letter to accuse OpenAI and Anthropic of Marxism
Aug 4, 2026
- 02
The Patch
The Patch — August 4, 2026
Aug 4, 2026
- 03
News
The rug pulled, twice — Timothy Gowers wrote the mathematical-culture case against Astra six days before OpenAI announced Astra
Aug 3, 2026
Letters
Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.