TL;DR Keep Opus 5.5 on
mediumfor everyday repo work. I gave it three tasks full of hidden edge cases (an open redirect, monthly renewals and a CSV importer) atmedium,highandxhigh, three runs each. All 27 runs passed every hidden case.highcost 36% more, andxhighcost three times as much and took three times as long. When a specific task calls for more, raise it for that session only:claude --effort high.
Opus 5.5 starts at medium, one level below Opus 5's high. In my Opus 5.5 benchmark, high matched medium on both results and cost, but those tasks were small enough for any model to solve. The real question was still open: where does medium fall short?
Anthropic offers a hint. Their post Spending your effort concludes that higher effort pays off on tasks with lots of hidden edge cases, because at higher levels Claude tests and verifies more. So I built tasks like that.
What I measured
Three tasks in a small PHP repo. No prompt lists the edge cases: each one states the goal, and the repo ships a visible test that a shallow fix makes pass. Every run is then scored against a hidden suite the model never saw.
| Task | What the prompt asks | Hidden suite |
|---|---|---|
| Open redirect | Only follow ?next= if the browser stays on https://app.example.com |
45 URLs (33 that take the user off the site, 12 legitimate), checked with Node against the WHATWG URL standard browsers implement |
| Renewals | Renew on the signup day, or the last day of shorter months, at the same local time | 20 cases: month ends, leap years, drift, DST, time zones |
| CSV | Read real exports from Excel and Google Sheets correctly (RFC 4180) | 23 cases: line breaks inside quotes, doubled quotes, backslashes, BOM, CRLF |
I first tried three tasks with a spec listing every rule. medium scored 100% on all three, so I dropped the spec.
Each run starts from a clean copy of the repo:
claude -p "<task prompt>" --model claude-opus-5-5 --effort xhigh \
--setting-sources project --strict-mcp-config \
--no-session-persistence --dangerously-skip-permissions \
--output-format json
No --effort for medium. I confirmed the actual level of each setup with the probe from the effort chain. The full benchmark (repo, prompts, hidden suites, reference solutions and all 27 runs with their diffs) is in this zip.
Results
Medians per run, plus the total of each level's 9 runs:
| Effort | Hidden cases | Cost | API time | Thinking | Total |
|---|---|---|---|---|---|
medium (Default) |
100% | $0.18 | 31 s | 506 tokens | $1.69 |
high |
100% | $0.22 | 42 s | 1,095 tokens | $2.30 |
xhigh |
100% | $0.38 | 95 s | 5,483 tokens | $4.99 |
Not one hidden case failed across 27 runs. The only thing the level changed was what I paid and how long I waited.
The open redirect shows it best. At medium, Opus 5.5 worked out on its own that PHP's parse_url() doesn't read URLs the way a browser does, and wrote a function that mimics the browser's parser: it strips tabs and newlines, treats \ as /, and drops everything before the last @. In one of the three runs it also wrote 26 test cases of its own. Median: under a minute and $0.24. At xhigh it landed on the same function, with the same 45 out of 45, in 4 to 6.5 minutes and for $0.90 to $1.20 a run.
Where the money goes
Opus 5.5 charges $20 per million output tokens, and thinking counts as output. Adding up each level's 9 runs:
medium |
high |
xhigh |
|
|---|---|---|---|
| Thinking | 7,709 | 21,521 | 92,176 |
| Total output | 32,165 | 50,509 | 140,214 |
| Cache reads | 1.14 M | 1.35 M | 2.35 M |
xhigh thinks 12 times as much as medium and writes over 4 times as much. It also takes more turns (a median of 7 against 5), and every turn rereads the conversation. On a subscription you never see the dollars, but you feel them in your five-hour limit and in the wait.
When raising it does pay off
My tasks ran from 16 seconds to 6.5 minutes of API time. Anthropic's are a different scale: Terminal-Bench 3.0 includes tasks like building an 8-bit game console in Verilog. That's where the gap shows up: on an HTML sanitizer meant to block JavaScript injection, Fable 5.1 went from 1 attempt in 5 at low to 5 in 5 at xhigh.
So high effort makes sense when you hand Claude a long job you won't review step by step: a security audit, a whole migration, something you launch and leave running. For a task you'll read when it's done, medium already does that work.
And raise it for that session only:
claude --effort high
Inside a session, per the docs, /effort opens the slider: pick the level and press s to apply it to the current session only. Enter, or typing /effort high, saves it as your default for that model, and from then on every small task pays for it.
My settings.json has no Opus 5.5 entry under modelSettings, so it runs at its medium. With these numbers, that's where it stays. For a single turn that needs more, there's ultrathink.
Official docs: Model configuration: Adjust effort level · Spending your effort
What this doesn't measure
- Long tasks. Nothing here ran past 6.5 minutes. There is a point where
highstarts paying for itself, but it sits above this task size. - Open-ended builds. Every task here had one right answer. According to Anthropic, on a "build me an app" prompt the level changes how much Claude decides on its own.
- Sample size. One model, three repetitions, one day.
Requirements
- Claude Code v2.1.284. Every figure comes from that version, on October 2, 2026, at API list price.