Free PDF guide

Stop overpaying in Claude Code

Model and effort are two dials, not one. What better actually means, the diagnostic when something fails, and where the bill comes from.

Are you a professional Web developer?

You join the Gravitas newsletter. Unsubscribe anytime.

Confirmed on another device, or session expired? Enter your email again and it unlocks straight away.

Stop overpaying in Claude Code

What is inside

Eight pages, one idea per page. No routing table, no "use Sonnet for the easy stuff".

  1. What "better" actually means: capability, time and money.
  2. Two dials, not one scale: who you send, and how much rope you give them.
  3. What effort really is: proactivity, not permission.
  4. Proof they are separate mechanisms: Haiku has no effort dial.
  5. The diagnostic in three questions, and why you climb one step at a time.
  6. Where the bill comes from: thinking is billed as output tokens.

The criterion, in one sentence

Pick the smallest model that seems reasonable, and go up one step only when you can prove the friction is not yours.

That is the whole thing. The rest of the guide is how to prove it.

The trap is the word "better". Most people read it as reasoning capability, and with that reading the answer is always the biggest model at maximum effort. But you spend three things, not one: capability, time and money. Better is whatever solves your task while spending the least of all three.

There is also a reason to climb one step at a time: you never climb back down. Once you are on the most expensive model for everything, going back is psychologically very hard.

What stays inside the PDF is the full diagnostic tree and the price ladder. Those are the two pages worth having at hand when something fails.

Frequently asked questions

Is raising effort the same as switching model?

No. The model decides how deep into the problem it can see; effort decides how much it does without being asked. Two independent dials, so four combinations rather than one scale. You can send the small model with the whole afternoon, or the big one in a hurry.

Does low effort stop Claude from editing files or running tests?

No. At low effort it edits files and runs tests if you ask. What it will not do is go beyond what you asked: fewer tool calls, no preamble, terser summaries. Effort governs initiative, not permissions.

Why does the bill jump so much just from raising effort?

Because thinking is billed as output tokens, the most expensive rate there is. Raising effort multiplies exactly the expensive half. Switching model or effort mid-conversation also invalidates the prompt cache, so run /compact before you switch.

Does this hold outside Claude?

The mechanism does; the names do not. The control surface differs across models and vendors: some take a number of reasoning tokens, others take a word from low to max. Haiku 4.5 does not accept the effort parameter at all and uses a token budget instead. The two-dial idea survives all of them.

Keep reading