TL;DR For a well-scoped task (a bug, a rename, a script, a refactor whose design is settled), start the session with
/model sonnet. Sonnet 5.5 passed 12 out of 12 for $0.88 in total against Opus 5.5's $1.91, with a median of 13 seconds of API time per task against 22. Keep Opus 5.5, your Default, for the ambiguous or long work. And pick up front: switching models mid-session breaks the cache, and the new model loses the old one's reasoning.
Sonnet 5.5 shipped yesterday, September 28, 2026, six days after Opus 5.5. Anthropic says it's over 30% faster than Sonnet 5 and costs up to 30% less per task, at the same per-token price.
In Claude Code your Default is still Opus 5.5. Sonnet 5.5 only runs when you ask for it. So the question is whether it's worth asking for, and when. I measured it with the same harness I used for Opus 5.5 against Fable 5.1.
What Anthropic claims
| Benchmark | Sonnet 5.5 | Opus 5.5 | Sonnet 5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | 10.3% |
| CursorBench 4.0 | 55.5% | 57.8% | 34.1% |
| OSWorld 2.1 (computer use) | 80.1% | 81.8% | 57.0% |
| GDPval-AA (knowledge work) | 1844 | 1846 | 1449 |
These are Anthropic's numbers, not mine (the Terminal-Bench Opus 5.5 score is at xhigh). Sonnet 5.5 lands within a point or two of Opus 5.5 on the rest, and ahead of it on Terminal-Bench. Anthropic still says Opus 5.5 is clearly stronger at complex, open-ended work that needs sustained judgment.
Per token, Sonnet 5.5 costs half of Opus 5.5 ($2/$10 against $4/$20 per MTok). Cache reads cost the same on both: $0.20.
What I measured
The same 11-file PHP repo and the same four tasks from the Opus 5.5 benchmark: fix a rounding bug, rename a class across 5 files, write a script that prices a JSONL of usage, and pull duplicated validation out of three controllers. Each has its own automated check. Every run starts from a clean copy and is launched like this:
claude -p "<task prompt>" --model claude-sonnet-5-5 \
--setting-sources project --strict-mcp-config \
--no-session-persistence --dangerously-skip-permissions \
--output-format json
--setting-sources project leaves out my user settings.json, so each model runs at its own default effort: medium on Sonnet 5.5 and Opus 5.5, high on Sonnet 5. I confirmed that first with the probe from the effort chain. The fourth contestant is Sonnet 5.5 with --effort high.
All three models ran on the same day and the same Claude Code version: I didn't reuse last week's numbers. The whole benchmark (repo, prompts, checks, harness, and all 48 runs with each one's diff) is in this zip.
Results
48 runs: 4 configurations × 4 tasks × 3 repetitions. Medians per run, plus the total across each configuration's 12:
| Model | Done | Cost | API time | Turns | Total |
|---|---|---|---|---|---|
Sonnet 5.5 (medium) |
12/12 | $0.074 | 13.4 s | 5 | $0.88 |
Sonnet 5.5 (high) |
12/12 | $0.074 | 14.3 s | 4.5 | $0.93 |
Opus 5.5 (medium) |
12/12 | $0.162 | 21.8 s | 5 | $1.91 |
Sonnet 5 (high) |
9/12 | $0.191 | 38.7 s | 13.5 | $2.34 |
Sonnet 5.5 was cheaper and faster than both Opus 5.5 and Sonnet 5 on every one of the four tasks. Against Opus 5.5: 54% less cost and 39% less time, with the same 12 passes. Against Sonnet 5: 62% less cost, twice what Anthropic promises.
Sonnet 5's three failures are real. Twice it rounded with round($total / 1000000 * 100) / 100, which gives 0.57 where 0.58 was due, and once it created the refactor's class but left the validation in all three controllers. Six days ago, on the same repo, it failed one in twelve. With three repetitions that gap is variance, not a regression.
Why half
Against Opus 5.5, two effects multiply:
- Rate. Opus 5.5's tokens billed at Sonnet 5.5 prices: $1.91 down to $1.12, 42% less. Not 50%, because cache reads cost the same on both.
- Tokens. Both finish in 5 turns and read almost the same from cache, but Sonnet 5.5 writes 43% less output. At the same rate, that's 21% less.
- Both together: $1.91 down to $0.88, 54% less (0.58 × 0.79).
Against Sonnet 5 the rate is identical, so the whole saving is tokens. Sonnet 5 needed a median of 13.5 turns, and every turn re-reads the context: across its runs it read 5.3 million tokens from cache, where Sonnet 5.5 read 1.5 million.
On speed, be precise about what's compared: Anthropic's "over 30% faster" is generation speed. What I measured is API time per task, which drops mostly because there are fewer turns.
Raising it to high bought nothing
Sonnet 5.5 on high did exactly what it did on medium: 12 out of 12 and the same $0.074 median. Just as with Opus 5.5, medium isn't holding it back on tasks this size. For one turn that needs more, there's ultrathink.
When Sonnet 5.5, when Opus 5.5, when Fable
| Model | Reach for it when |
|---|---|
| Sonnet 5.5 | You know what has to change and how to check it: contained bugs, renames, scripts, refactors with the design settled, iterating on a feature |
| Opus 5.5 (Default) | You need to decide the what, not only the how: design, bugs that cross modules, long sessions with chained decisions |
| Fable 5.1 | Genuinely new territory, or a task bigger than one sitting |
In July I wrote that Sonnet lived inside subagents and never as the main model. With these numbers, for a well-scoped task, that no longer holds.
The first row is what this benchmark measures. The other two it doesn't: here both models pass everything, and Opus 5.5's edge on judgment-heavy work is Anthropic's claim. For Fable 5.1, the reasoning is in how I route the models.
Two things change without you touching anything. Your subagents with model: sonnet now run on Sonnet 5.5. And opusplan, which plans with Opus and executes with Sonnet, now executes with Sonnet 5.5.
Pick at the start of the session
Switching models mid-session costs you twice:
- The cache. The model is part of the cache key. The turn after the switch re-reads the whole conversation at full price (the detail).
- The reasoning. No other model reads Sonnet 5.5's thinking blocks, and Sonnet 5.5 can't read Opus 5.5's. Switch between them and the API drops those blocks, so the new model carries on without that reasoning.
So decide when you open the session:
claude --model sonnet
Or /model sonnet before the first prompt. If the task turns harder along the way, /compact and then /model opus: the switch lands on a history that's already short.
What changes in Claude Code
- It needs v2.1.284, the release that adds it and points the
sonnetalias at it. After updating,claude --model sonnetanswers asclaude-sonnet-5-5on my account. - It starts at
medium, and your global effort doesn't reach it. My usersettings.jsonhas"effortLevel": "high", and the probe returnsEFFORT=mediumon Sonnet 5.5. Same as Opus 5.5: for another level, addclaude-sonnet-5-5undermodelSettings. - Thinking can't be turned off.
alwaysThinkingEnabledandMAX_THINKING_TOKENS=0do nothing; effort is the only control (the rest, in thinking). - No fast mode.
/fastexists only on Opus: turning it on from Sonnet 5.5 moves you to Opus. - Native 1M context. The
[1m]suffix does nothing on Sonnet 5.5. - Refusals. When a classifier flags a request for cybersecurity, Claude Code re-runs it on Sonnet 5. A biology flag ends in a refusal.
Official docs: Model configuration · Introducing Claude Sonnet 5.5 · What's new in Claude Sonnet 5.5
Coming from Sonnet 5? What its launch brought is in Sonnet 5 in Claude Code. And why Opus 5.5 is your Default, in Opus 5.5 in Claude Code.
What this doesn't measure
- Quality on hard tasks. Sonnet 5.5 and Opus 5.5 pass everything, so this separates cost and speed, not reasoning.
- Scale. Tasks of 7 to 50 seconds of API time, three repetitions, one day.
- API dollars. On a subscription you never see this figure. What you notice is how long your five-hour limit lasts and how long you wait.
Requirements
- Claude Code v2.1.284 or later (
claude update). Every figure comes from that version, on September 29, 2026.