← Claude Code Hub
✦ Tip #213 Sep 29, 2026

Sonnet 5.5 vs Opus 5.5 in Claude Code: the same work at half the cost, and faster

Anthropic says Sonnet 5.5 costs up to 30% less than Sonnet 5. I tested it over 48 runs in Claude Code, against Sonnet 5 and against Opus 5.5, your default: it passed everything for under half of what Opus 5.5 cost. Plus when Opus is still worth it.

Sonnet 5.5 vs Opus 5.5 vs Sonnet 5: the same 4 tasks in Claude Code, 3 runs each. Total cost of 12 runs: Sonnet 5.5 $0.88, Opus 5.5 $1.91, Sonnet 5 $2.34. Median API time: 13.4, 21.8 and 38.7 seconds

TL;DR For a well-scoped task (a bug, a rename, a script, a refactor whose design is settled), start the session with /model sonnet. Sonnet 5.5 passed 12 out of 12 for $0.88 in total against Opus 5.5's $1.91, with a median of 13 seconds of API time per task against 22. Keep Opus 5.5, your Default, for the ambiguous or long work. And pick up front: switching models mid-session breaks the cache, and the new model loses the old one's reasoning.

Sonnet 5.5 shipped yesterday, September 28, 2026, six days after Opus 5.5. Anthropic says it's over 30% faster than Sonnet 5 and costs up to 30% less per task, at the same per-token price.

In Claude Code your Default is still Opus 5.5. Sonnet 5.5 only runs when you ask for it. So the question is whether it's worth asking for, and when. I measured it with the same harness I used for Opus 5.5 against Fable 5.1.

What Anthropic claims

Benchmark Sonnet 5.5 Opus 5.5 Sonnet 5
Terminal-Bench 4.0 70.6% 66.4% 10.3%
CursorBench 4.0 55.5% 57.8% 34.1%
OSWorld 2.1 (computer use) 80.1% 81.8% 57.0%
GDPval-AA (knowledge work) 1844 1846 1449

These are Anthropic's numbers, not mine (the Terminal-Bench Opus 5.5 score is at xhigh). Sonnet 5.5 lands within a point or two of Opus 5.5 on the rest, and ahead of it on Terminal-Bench. Anthropic still says Opus 5.5 is clearly stronger at complex, open-ended work that needs sustained judgment.

Per token, Sonnet 5.5 costs half of Opus 5.5 ($2/$10 against $4/$20 per MTok). Cache reads cost the same on both: $0.20.

What I measured

The same 11-file PHP repo and the same four tasks from the Opus 5.5 benchmark: fix a rounding bug, rename a class across 5 files, write a script that prices a JSONL of usage, and pull duplicated validation out of three controllers. Each has its own automated check. Every run starts from a clean copy and is launched like this:

claude -p "<task prompt>" --model claude-sonnet-5-5 \
  --setting-sources project --strict-mcp-config \
  --no-session-persistence --dangerously-skip-permissions \
  --output-format json

--setting-sources project leaves out my user settings.json, so each model runs at its own default effort: medium on Sonnet 5.5 and Opus 5.5, high on Sonnet 5. I confirmed that first with the probe from the effort chain. The fourth contestant is Sonnet 5.5 with --effort high.

All three models ran on the same day and the same Claude Code version: I didn't reuse last week's numbers. The whole benchmark (repo, prompts, checks, harness, and all 48 runs with each one's diff) is in this zip.

Results

48 runs: 4 configurations × 4 tasks × 3 repetitions. Medians per run, plus the total across each configuration's 12:

Model Done Cost API time Turns Total
Sonnet 5.5 (medium) 12/12 $0.074 13.4 s 5 $0.88
Sonnet 5.5 (high) 12/12 $0.074 14.3 s 4.5 $0.93
Opus 5.5 (medium) 12/12 $0.162 21.8 s 5 $1.91
Sonnet 5 (high) 9/12 $0.191 38.7 s 13.5 $2.34

Sonnet 5.5 was cheaper and faster than both Opus 5.5 and Sonnet 5 on every one of the four tasks. Against Opus 5.5: 54% less cost and 39% less time, with the same 12 passes. Against Sonnet 5: 62% less cost, twice what Anthropic promises.

Sonnet 5's three failures are real. Twice it rounded with round($total / 1000000 * 100) / 100, which gives 0.57 where 0.58 was due, and once it created the refactor's class but left the validation in all three controllers. Six days ago, on the same repo, it failed one in twelve. With three repetitions that gap is variance, not a regression.

Why half

Against Opus 5.5, two effects multiply:

  • Rate. Opus 5.5's tokens billed at Sonnet 5.5 prices: $1.91 down to $1.12, 42% less. Not 50%, because cache reads cost the same on both.
  • Tokens. Both finish in 5 turns and read almost the same from cache, but Sonnet 5.5 writes 43% less output. At the same rate, that's 21% less.
  • Both together: $1.91 down to $0.88, 54% less (0.58 × 0.79).

Against Sonnet 5 the rate is identical, so the whole saving is tokens. Sonnet 5 needed a median of 13.5 turns, and every turn re-reads the context: across its runs it read 5.3 million tokens from cache, where Sonnet 5.5 read 1.5 million.

On speed, be precise about what's compared: Anthropic's "over 30% faster" is generation speed. What I measured is API time per task, which drops mostly because there are fewer turns.

Raising it to high bought nothing

Sonnet 5.5 on high did exactly what it did on medium: 12 out of 12 and the same $0.074 median. Just as with Opus 5.5, medium isn't holding it back on tasks this size. For one turn that needs more, there's ultrathink.

When Sonnet 5.5, when Opus 5.5, when Fable

Model Reach for it when
Sonnet 5.5 You know what has to change and how to check it: contained bugs, renames, scripts, refactors with the design settled, iterating on a feature
Opus 5.5 (Default) You need to decide the what, not only the how: design, bugs that cross modules, long sessions with chained decisions
Fable 5.1 Genuinely new territory, or a task bigger than one sitting

In July I wrote that Sonnet lived inside subagents and never as the main model. With these numbers, for a well-scoped task, that no longer holds.

The first row is what this benchmark measures. The other two it doesn't: here both models pass everything, and Opus 5.5's edge on judgment-heavy work is Anthropic's claim. For Fable 5.1, the reasoning is in how I route the models.

Two things change without you touching anything. Your subagents with model: sonnet now run on Sonnet 5.5. And opusplan, which plans with Opus and executes with Sonnet, now executes with Sonnet 5.5.

Pick at the start of the session

Switching models mid-session costs you twice:

  • The cache. The model is part of the cache key. The turn after the switch re-reads the whole conversation at full price (the detail).
  • The reasoning. No other model reads Sonnet 5.5's thinking blocks, and Sonnet 5.5 can't read Opus 5.5's. Switch between them and the API drops those blocks, so the new model carries on without that reasoning.

So decide when you open the session:

claude --model sonnet

Or /model sonnet before the first prompt. If the task turns harder along the way, /compact and then /model opus: the switch lands on a history that's already short.

What changes in Claude Code

  • It needs v2.1.284, the release that adds it and points the sonnet alias at it. After updating, claude --model sonnet answers as claude-sonnet-5-5 on my account.
  • It starts at medium, and your global effort doesn't reach it. My user settings.json has "effortLevel": "high", and the probe returns EFFORT=medium on Sonnet 5.5. Same as Opus 5.5: for another level, add claude-sonnet-5-5 under modelSettings.
  • Thinking can't be turned off. alwaysThinkingEnabled and MAX_THINKING_TOKENS=0 do nothing; effort is the only control (the rest, in thinking).
  • No fast mode. /fast exists only on Opus: turning it on from Sonnet 5.5 moves you to Opus.
  • Native 1M context. The [1m] suffix does nothing on Sonnet 5.5.
  • Refusals. When a classifier flags a request for cybersecurity, Claude Code re-runs it on Sonnet 5. A biology flag ends in a refusal.

Official docs: Model configuration · Introducing Claude Sonnet 5.5 · What's new in Claude Sonnet 5.5

Coming from Sonnet 5? What its launch brought is in Sonnet 5 in Claude Code. And why Opus 5.5 is your Default, in Opus 5.5 in Claude Code.

What this doesn't measure

  • Quality on hard tasks. Sonnet 5.5 and Opus 5.5 pass everything, so this separates cost and speed, not reasoning.
  • Scale. Tasks of 7 to 50 seconds of API time, three repetitions, one day.
  • API dollars. On a subscription you never see this figure. What you notice is how long your five-hour limit lasts and how long you wait.

Requirements

  • Claude Code v2.1.284 or later (claude update). Every figure comes from that version, on September 29, 2026.
Free guide

The 51 essentials, as a guide.

One page per tip. Five chapters. What I actually use daily in production. No theory, no fluff.

  • I. Getting started 10 tips
  • II. Awareness 3 tips
  • III. Mastery 22 tips
  • IV. Autonomy 10 tips
  • V. Comparison 6 tips
Are you a professional Web developer?

You'll receive the guide by email · You join the Gravitas newsletter · Unsubscribe anytime

of 51
#

Wmedia · 51 Tips
Free guide · 51 tips · 5 chapters

The 51 essentials, as a guide.

Are you a professional Web developer? · Unsubscribe anytime
Workshop for teams

Multiply your team's output without sacrificing quality: a 6 to 8 hour AI First workshop, online, on the Claude platform.

See the workshop

Want the 51 Claude Code essentials as a guide?