TL;DR Give your repeat work its own subagent with
model: haikuand only the tools it needs. Mine starts at 8K tokens per request and pays Haiku 5.5's low rate, 90% under Haiku 4.5. My main session on Haiku starts at 51K, already halfway to the 100K line where the discount drops to 50%.
Claude Haiku 5.5 (claude-haiku-5-5) shipped on October 7, 2026, and from v2.1.293 it's what answers when Claude Code asks for haiku on the Anthropic API. It costs $0.10/$0.50 per MTok against Haiku 4.5's $1/$5, but only while the prompt stays under 100K tokens. Above that, it's $0.50/$2.50. In Claude Code, that line decides where you save 90% and where you only save half.
What's new
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 70.6% |
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 83.9% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | 64.5% |
These are Anthropic's numbers, not mine. It doesn't reach Sonnet 5.5, but the jump from Haiku 4.5 is huge, agentic coding most of all. It's also the first Haiku with adjustable effort (low to max, medium by default), it has a 1M context window, and Anthropic calls it their fastest model yet. They pitch it "as a subagent on coding work", next to Opus 5.5 and Sonnet 5.5.
The price has a step at 100K
| Per MTok | Haiku 5.5 (up to 100K) | Haiku 5.5 (over 100K) | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1 |
| Output | $0.50 | $2.50 | $5 |
| Cache write (1 h) | $0.20 | $1 | $2 |
| Cache read | $0.01 | $0.05 | $0.10 |
The step is applied per request, by prompt length. And every Claude Code request carries the system prompt, the tool definitions, your CLAUDE.md and the whole conversation so far.
How big a Claude Code request is
I measured it in this repo on v2.1.293. First, the main session on Haiku with a one-word question:
claude -p --model haiku --output-format json "Reply with the single word OK"
Then a second, identical session, this time handing a search to the finder subagent shown below: Haiku and three tools. Per-request sizes, read from the transcripts:
main session (Haiku 5.5) 51,059 tokens per request
finder subagent (Haiku 5.5) 8,097 tokens per request
The main session is halfway to the line before it does anything. Read a handful of files and the conversation goes past 100K, onto the high rate. That's still half of Haiku 4.5's price, but no longer a tenth.
The subagent starts with a context of its own, without your conversation, and with only the tools you gave it. There's room for a lot of work before it gets near 100K.
One more thing: Haiku 5.5 counts more tokens than Haiku 4.5 for the same text. The same question from the same folder came to 47,311 tokens on Haiku 5.5 and 40,032 on Haiku 4.5, 18% more. A subagent that sat around 85K on Haiku 4.5 can cross 100K on Haiku 5.5.
Where it pays off: your own Haiku subagent
1. Create the subagent
In .claude/agents/finder.md (or ~/.claude/agents/ for every project):
---
name: finder
description: Finds the files in the repo that mention a term and reports their paths
tools: Glob, Grep, Read
model: haiku
---
You search the repository with Grep and Glob and report the matching file paths, one per line.
model: haiku is what takes it off your main model. Without that line, the subagent runs on your session's model. tools is what keeps it light: every tool you leave out is one less definition in every request. The rest of the format is in creating custom agents.
Good candidates are searches, summarizing logs or docs, classifying or extracting data: mechanical work you repeat, where speed matters more than reasoning.
2. Check which model answered
Ask Claude to use the subagent and run the test with -p --output-format json: the modelUsage field lists claude-haiku-5-5 with its tokens. In an interactive session, /usage breaks it down under "Usage by model".
To send every subagent that has no model of its own to Haiku, there's CLAUDE_CODE_SUBAGENT_MODEL, which I covered in effort per model. I'd rather set it in each definition: the general-purpose agent picks up that variable too, and that one does work I want on Opus.
What already runs on Haiku without you doing anything
Since v2.1.293, all of this runs on Haiku 5.5 on the Anthropic API:
- The
claude-code-guidesubagent, the one that answers when you ask how Claude Code works. - Prompt hooks, unless you give them a
model. - The
/goalevaluator, which decides after each turn whether the condition holds. - Background jobs, such as the summaries
claude --resumeuses.
ANTHROPIC_DEFAULT_HAIKU_MODEL controls all four: it changes what haiku resolves to and what model they run on.
What doesn't run on Haiku is the Explore agent. It runs on your main session's model (on Opus when you're on Fable). If you read otherwise in my model table, that row is out of date.
The fine print
- Requires Claude Code v2.1.293 (
claude update). That's the version the docs ask for to use Haiku 5.5. - Anthropic API only. On Bedrock, Google Cloud's Agent Platform, Microsoft Foundry and Claude Platform on AWS,
haikustill resolves to Haiku 4.5. - A session saved on Haiku 4.5 resumes on Haiku 5.5 if your
modelsetting ishaiku. - Thinking can't be turned off. As on Opus 5.5 and Sonnet 5.5, effort is the only control.
Official docs: Model configuration · Subagents · Introducing Claude Haiku 5.5
Requirements
- Claude Code v2.1.293 or later. Every measurement above comes from that version, on this repo's account on the Anthropic API. Your exact request size depends on your
CLAUDE.md, MCP servers and skills: measure yours with the same command.