
Claude Haiku 5.5 launched October 7, 2026, with a 1M context window and up to 90% lower pricing than Haiku 4.5. Here's what's confirmed, what it costs, and the token-usage catch.
Sai Meghana G
Software Engineer
Claude Haiku 5.5 isn't coming soon. It's already here. Anthropic shipped it on October 7, 2026, with a 1 million token context window, pricing as low as $0.10 per million input tokens, and benchmark jumps that make Haiku 4.5 look like a different product line. If you'd started reading this expecting a "still waiting" story, the story moved while nobody was looking. That's worth sitting with for a second: the leak-to-launch window for this entire model generation ran about two weeks.
Here's what's actually confirmed, what changed from Haiku 4.5, and the one catch that's easy to miss in the price-cut headlines.
It was, briefly. Here's the actual timeline:
| Date | What happened |
|---|---|
| Sept 22, 2026 | Anthropic launches Claude Opus 5.5 and confirms, in the same post, that "Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks" |
| Sept 28, 2026 | Claude Sonnet 5.5 launches, keeping Sonnet 5's pricing but running about 30% faster |
| Oct 7, 2026 | Claude Haiku 5.5 launches, generally available on Claude.ai and the API |
Two models in six days, then Haiku nine days after that. Anthropic moved through its entire 5.5 lineup in under three weeks. If you saw an earlier draft of this story treating Haiku 5.5 as speculative, that's already stale — which is exactly the kind of thing that happens when you cover a model a company hasn't shipped yet.
Anthropic ships Claude in a few sizes. Opus does the heavy thinking. Sonnet is the all-rounder most people default to. Haiku is the one built to run thousands or millions of times a day without wrecking your bill — support chats, email sorting, document summaries, subagents that do one small job inside a bigger workflow.
Haiku was never going to win a benchmark fight with Opus. That's not its job. Its job is to be cheap and fast enough that you stop thinking about the per-call cost at all.
This is the headline Anthropic wants you to remember, and it holds up:
| Haiku 4.5 | Haiku 5.5 (≤100K tokens) | Haiku 5.5 (>100K tokens) | |
|---|---|---|---|
| Input | $1 / 1M tokens | $0.10 / 1M tokens | $0.50 / 1M tokens |
| Output | $5 / 1M tokens | $0.50 / 1M tokens | $2.50 / 1M tokens |
| Cache reads | — | $0.01 / 1M tokens | $0.05 / 1M tokens |
For a typical prompt under 100,000 tokens, that's a 90% cut on input and output pricing. Anthropic's own framing lands at "around 75% less on average," which accounts for the fact that longer prompts fall into the pricier tier and that Haiku 5.5 uses a different tokenizer (more on that below). Either number you use, this is a real price cut, not a rounding trick.
Haiku 5.5 supports a 1 million token context window and up to 128,000 output tokens. That matters more than it sounds: Haiku 4.5 topped out well short of that, so this is the first time Anthropic's cheapest model has matched the context size of its flagship Opus 5.5 and Sonnet 5.5 siblings. You can now feed a Haiku-tier model an entire codebase or a long document chain without hitting a wall that used to force you up to Sonnet.
Anthropic's own release numbers show the biggest generation-over-generation jump I've seen for a Haiku release:
| Benchmark | Haiku 4.5 | Haiku 5.5 |
|---|---|---|
| OSWorld 2.1 (computer use) | 15.7% | 72.4% |
| Terminal-Bench 4.0 (agentic coding) | 0.0% | 39.2% |
| Humanity's Last Exam (with tools) | 18.7% | 57.4% |
| GDPval-AA v2.1 | 735 | 1,620 |
| AA-Briefcase v1.1 | 614 | 1,578 |
| Chartography | 6.4% | 46.4% |
Terminal-Bench going from a flat zero to 39.2% is the number that stands out to me. Haiku 4.5 basically couldn't do agentic terminal work. Haiku 5.5 can, at least often enough to be useful. Third-party tracking backs this up too: Artificial Analysis put Haiku 5.5 at 43 on its Intelligence Index, up 26 points from the last Haiku release a year earlier.
Haiku 5.5 is also the first Haiku model with an adjustable effort setting, the same "how hard should this model think" dial Anthropic added to Opus and Sonnet. Low effort for quick, cheap calls; higher effort when you need Haiku to actually reason through something instead of pattern-matching its way to an answer.
Here's the part that doesn't make the pricing slide. Haiku 5.5 runs on an updated tokenizer, the same change Anthropic made for Sonnet 5.5 and Opus 5.5, and it uses more tokens to express the same amount of work than Haiku 4.5 did. Anthropic says its pricing comparisons already account for this shift, which is a fair way to handle it.
But "more tokens per task" shows up again when you compare Haiku 5.5 to competitors. Artificial Analysis found that at matched effort settings, Haiku 5.5 writes something like 3x the output tokens GPT-6 Luna uses to land a similar score, and at max effort it burns around 162,000 output tokens per Intelligence Index task, more than Opus 5.5 uses on the same benchmark. A cheap per-token rate doesn't automatically mean a cheap bill if the model needs three times as many tokens to finish the job. Test this against your actual workload before you assume the sticker price is the real price.
The two models share identical headline pricing below 100,000 tokens: $0.10 per million input, $0.50 per million output, on both sides. Where they diverge is the pricing cliff. Haiku 5.5's rate jumps at 100,000 tokens; Luna's doesn't move until 272,000 tokens. Run a prompt in that 100K–272K range and Luna comes out roughly 5x cheaper for that slice of work.
On capability, multiple independent write-ups report Haiku 5.5 beating GPT-6 Luna on every benchmark where both have a published score in Anthropic's release chart. Just remember the token-spend asterisk above before you call that a clean win on cost.
Haiku 5.5 rolled out wide on day one:
claude-haiku-5-5Asana is already running it in production. Aaron Vinh, an engineer there, said the switch brought "over 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn." That tracks with what Haiku is actually for: not the smartest model in the room, but the one you can afford to call constantly.
No, and this is worth closing the loop on because it was a real story for about a week. In mid-September, social posts claimed Anthropic was retiring the Haiku line entirely, with only Fable, Opus, and Sonnet getting 5.5 versions. Anthropic's own Opus 5.5 launch post killed that rumor directly, naming Haiku 5.5 by name and promising it within weeks. Nine days later, it shipped.
One naming quirk worth flagging: there was never a Haiku 5. Anthropic jumped straight from Haiku 4.5 to Haiku 5.5, matching the version number Opus and Sonnet landed on. If you go looking for "Claude Haiku 5" expecting a separate release, you won't find one.
My honest read: yes, almost certainly, if you're running high-volume, low-complexity work. The benchmark gains are large enough that you're not trading capability for price here, you're getting both. A 90% price cut on short prompts and a jump from 0% to 39.2% on agentic terminal tasks is not a subtle upgrade.
The one place I'd pump the brakes: if you've tuned a workload tightly around Haiku 4.5's token economics, don't assume the new tokenizer and the longer effort settings will behave identically. Run your own cost test on a slice of real traffic before you flip production traffic over. "Cheaper per token" and "cheaper per task" aren't always the same number, and Haiku 5.5 is a good example of why that gap matters.
For the rest of the Claude 5.5 lineup, our Claude Opus 5 pricing breakdown and Opus 5 vs Fable 5 comparison cover where the bigger models land on cost and benchmarks. And if you're weighing Anthropic against Google's side of this race, Gemini 3.8 Flash is the closest competing workhorse tier.
Three models, three weeks, and a rumor that didn't survive the first one. That's the actual Haiku 5.5 story: not a wait, but a sprint.
New AI tools, automation workflows, and course drops — straight to your inbox. Join 2,400+ builders.

Google launched Nano Banana 2.1 on October 6, 2026, cutting image API prices in half to $0.0336 per 1K image while improving text rendering, editing, and character consistency. Here's what actually changed.

Gemini 3.8 Flash launched Sept 2, 2026 at $0.75/$3.75 per 1M tokens, with a 1M context window. Here's what's new, the hidden token costs, and who should actually switch.

A solo dev built PaperRoute, a finished Paperboy-style browser game, with GPT-6 Astra, Blender and Meshy in 39 tracked hours and 1.56B tokens. Here's the 6-step process he used.