{"id":772,"date":"2026-10-10T05:04:25","date_gmt":"2026-10-10T05:04:25","guid":{"rendered":"https:\/\/abrarqasim.com\/blog\/claude-haiku-5-5-pricing-the-1-25x-token-count-behind-the-price-tag\/"},"modified":"2026-10-10T05:04:25","modified_gmt":"2026-10-10T05:04:25","slug":"claude-haiku-5-5-pricing-the-1-25x-token-count-behind-the-price-tag","status":"publish","type":"post","link":"https:\/\/abrarqasim.com\/blog\/claude-haiku-5-5-pricing-the-1-25x-token-count-behind-the-price-tag\/","title":{"rendered":"Claude Haiku 5.5 Pricing: The Token Count Behind the Price Tag"},"content":{"rendered":"<p>Okay, this is going to sound pedantic, but the first thing I did after reading about Claude Haiku 5.5 was ignore the price. I looked at the tokenizer line instead. A new Haiku that costs $0.10 per million input tokens sounds like a clean 10x discount on the old one, and I&rsquo;ve learned to distrust clean discounts. The per-token price is only half of a bill. The other half is how many tokens your prompt turns into.<\/p>\n<p>Short version for the impatient: Haiku 5.5 is much cheaper than Haiku 4.5 for small requests, roughly the same price as OpenAI&rsquo;s GPT-6 Luna under 100,000 tokens, and a worse deal above that. The tokenizer change eats about a fifth of the savings before you notice. Everything below comes from <a href=\"https:\/\/simonwillison.net\/2026\/Oct\/7\/claude-haiku-5-5\/\" rel=\"nofollow noopener\" target=\"_blank\">Simon Willison&rsquo;s write-up<\/a>, who measured it, plus my own arithmetic. I haven&rsquo;t run Haiku 5.5 on a client workload yet, so treat the numbers as a worked example and not a benchmark.<\/p>\n<h2 id=\"the-price-tag-as-published\">The price tag, as published<\/h2>\n<p>Simon&rsquo;s post gives the numbers. Haiku 4.5 was priced at $1 per million input tokens and $5 per million output, which he calls relatively expensive even when it launched. Haiku 5.5 is $0.10 and $0.50, matching GPT-6 Luna, which came out last month. That&rsquo;s a tenth of the old price on both sides.<\/p>\n<p>Then the footnotes start. Past 100,000 tokens, the Haiku 5.5 price goes up five times, to $0.50 and $2.50. Luna also has a step, but it kicks in at 272,000 tokens and only goes to $0.20 and $0.75. So the headline price is real, but it has a cliff, and the cliff sits in a place where plenty of real prompts live. A long document, a big retrieval dump, or an agent that keeps appending tool results to its history can all cross 100,000 tokens without anyone noticing.<\/p>\n<p>I&rsquo;m not sure from the post whether the higher price applies to the whole request once you cross the line or only to the tokens past it. That changes the math a lot, so check the <a href=\"https:\/\/docs.claude.com\/en\/docs\/about-claude\/pricing\" rel=\"nofollow noopener\" target=\"_blank\">pricing docs<\/a> before you budget. I&rsquo;ll show the whole-request case below because it&rsquo;s the pessimistic one.<\/p>\n<h2 id=\"the-tokenizer-line-that-matters-more\">The tokenizer line that matters more<\/h2>\n<p>Simon also notes that Haiku 5.5 uses a new tokenizer. He ran the same long prompt through his token counter and got about 1.25 times as many tokens with Haiku 5.5 as with Haiku 4.5. His phrase for it is a hidden price increase, and I agree with him.<\/p>\n<p>Here&rsquo;s why it bites. Your bill is price per token times token count. If a prompt that was 10,000 tokens becomes 12,500, then a 10x price cut gives you an 8x cut. Still great. But if you were comparing against another vendor by looking only at the rate card, you were comparing the wrong numbers. And the 100,000 token cliff moves too: a prompt that was comfortably under it on the old tokenizer can land past it on the new one. That is the case that worries me.<\/p>\n<p>This is also where the token counter becomes an actual tool instead of a toy. The only way to know your real count is to measure your own prompts with the new model&rsquo;s tokenizer. A rule of thumb like &ldquo;about four characters per token&rdquo; tells you very little once the tokenizer changes under you.<\/p>\n<h2 id=\"a-worked-example-with-my-own-numbers\">A worked example with my own numbers<\/h2>\n<p>Say you have a classification job in a support tool. Each request carries a 10,000 token prompt on the old tokenizer and returns 500 tokens. I&rsquo;m assuming the output also inflates by the same 1.25 factor, which the post doesn&rsquo;t say, so that part is my guess.<\/p>\n<pre><code class=\"language-python\">def cost(input_tokens, output_tokens, in_price, out_price):\n    return (input_tokens * in_price + output_tokens * out_price) \/ 1_000_000\n\n# Haiku 4.5: $1 in, $5 out per million tokens\nold = cost(10_000, 500, 1.00, 5.00)\n\n# Haiku 5.5: $0.10 in, $0.50 out, 1.25x tokens, under the 100k step\nnew = cost(12_500, 625, 0.10, 0.50)\n\nprint(f&quot;old ${old:.5f}  new ${new:.5f}  ratio {old \/ new:.1f}x&quot;)\n# old $0.01250  new $0.00156  ratio 8.0x\n<\/code><\/pre>\n<p>An 8x cut, which is nothing to complain about. Across a million of those requests a month, that&rsquo;s $12,500 versus about $1,560 by my arithmetic. Good news, and I&rsquo;d take it.<\/p>\n<p>Now the long-prompt case. A document question with a 120,000 token prompt on the old tokenizer, which is 150,000 on the new one. If the higher price applies to the whole request, the input alone costs 150,000 times $0.50 per million, so $0.075. On Haiku 4.5 the same input cost $0.12. You still save money, but it&rsquo;s a 37 percent saving instead of a 90 percent one. If your workload lives up there, &ldquo;Haiku is 10x cheaper&rdquo; is not a sentence you can say to a client.<\/p>\n<p>One more caveat. Simon says that above 100,000 tokens Luna looks like a much better deal, but token counts differ between vendors, so I wouldn&rsquo;t compare the two at a fixed count. Count the same prompt with both tokenizers first.<\/p>\n<h2 id=\"the-other-changes-in-the-announcement\">The other changes in the announcement<\/h2>\n<p>Two other things came with the launch. Anthropic halved the price of cache reads for Sonnet 5.5, which matters if your agent re-sends a big stable prefix on every turn. I covered the arithmetic for that style of workload in my <a href=\"https:\/\/abrarqasim.com\/blog\/llm-pricing-comparison-cached-token-math-opus-5-5-gpt-6-grok\" rel=\"noopener\">LLM pricing comparison post<\/a>, and the cache-read discount only makes the cached path look better.<\/p>\n<p>The second is API credits for subscription plans. Per the quote Simon includes, Max 5x subscribers get $100 a month, Max 20x get $200, and Team subscribers get up to $500, pooled. They don&rsquo;t roll over. You can also turn off auto-reload, so requests stop when your balance runs out instead of billing you further. For a freelancer experimenting with prototypes, that spend cap is the part I&rsquo;d use first. I&rsquo;ve seen enough surprise invoices from a runaway loop that a hard stop feels like a feature.<\/p>\n<p>Simon makes a related point in a post from a few days earlier, that we&rsquo;re going to need default hard budget caps on pretty much everything. I agree. The honest takeaway is that cost control in LLM work is mostly about limits you set yourself, not prices you read.<\/p>\n<h2 id=\"what-id-do-before-switching-a-production-job\">What I&rsquo;d do before switching a production job<\/h2>\n<p>I would not move a working job to Haiku 5.5 because the rate card dropped. I&rsquo;d do these in order:<\/p>\n<ol>\n<li>Take 50 real prompts from your logs, run them through the token counter for both models, and record the ratio. Don&rsquo;t trust 1.25 for your data. It&rsquo;s one prompt on one tool.<\/li>\n<li>Find out how many of those prompts sit between 80,000 and 100,000 tokens on the old tokenizer, because those are the ones the new tokenizer may push past the step.<\/li>\n<li>Re-run your quality checks. Simon says the benchmark scores are higher, but a general benchmark says nothing about your classification labels or your extraction schema.<\/li>\n<li>Set a budget cap on the API key before you flip the switch.<\/li>\n<\/ol>\n<p>For the quality check, I like grading outputs against a small frozen set of cases, the same habit I use for agents. If you want a method, I wrote about <a href=\"https:\/\/abrarqasim.com\/blog\/ai-agent-evaluation-grade-the-database-not-the-reply\" rel=\"noopener\">grading the database instead of the reply<\/a>, and the same idea works for any extraction job.<\/p>\n<p>If you want help auditing an LLM bill for a product that&rsquo;s already live, that&rsquo;s work I do, and you can see examples on my <a href=\"https:\/\/abrarqasim.com\" rel=\"noopener\">portfolio<\/a>. Before you hire anyone, though, run the first step above yourself. It takes an afternoon.<\/p>\n<p>This week, pull 50 production prompts, count them with the Haiku 5.5 tokenizer, and write the ratio and the number above 100,000 in a note next to your current monthly bill. If the ratio is near 1.25 and nothing crosses the step, the switch is probably a win. If anything crosses, price that slice separately.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Haiku 5.5 costs a tenth of Haiku 4.5 per token, but a new tokenizer and a 100k step change the bill. A worked example and what I would measure before switching.<\/p>\n","protected":false},"author":2,"featured_media":771,"comment_status":"","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"","rank_math_description":"Haiku 5.5 costs a tenth of Haiku 4.5 per token, but a new tokenizer and a 100k step change the bill. A worked example and what I would measure before switching.","rank_math_focus_keyword":"claude token counter","rank_math_canonical_url":"","rank_math_robots":"","footnotes":""},"categories":[4],"tags":[521,467,851,668,852],"class_list":["post-772","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai","tag-anthropic","tag-claude","tag-haiku","tag-llm-pricing","tag-tokenizer"],"_links":{"self":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/posts\/772","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/comments?post=772"}],"version-history":[{"count":0,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/posts\/772\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/media\/771"}],"wp:attachment":[{"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/media?parent=772"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/categories?post=772"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/abrarqasim.com\/blog\/wp-json\/wp\/v2\/tags?post=772"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}