Low, medium, high or max: how much thinking should you pay an AI model for?
Claude and GPT-6 both let you choose how hard the model thinks, from low to max, and the thinking is billed as output. Start at medium, test low on real examples, and raise the level only for the tasks that fail.
Anthropic posted a short guide this week on choosing an effort level for Claude Opus 5.5. Effort is the setting that decides how hard the model thinks before it answers, and most people never touch it. OpenAI has the same dial on its GPT-6 models, where it is called reasoning effort.
You should care about this setting because you end up paying for it. OpenAI says the reasoning tokens are billed as output tokens even though you never see them, and on Claude the effort level shapes every token in the answer, including the thinking and the tool calls.
Here is where each model starts if you never set it, straight from each lab's documentation.
At the top, Claude Fable 5.1 (and the invitation-only Mythos 5.1) starts at high, and GPT-6 Astra doesn't publish a default. Just below, Claude Opus 5.5 starts at medium. In the middle, Claude Sonnet 5 starts at high and GPT-6 Sol at medium. At the small end, Claude Haiku 4.5 has no effort setting, while GPT-6 Luna starts at medium.
Every model with the setting goes up to max, and only Sol and Luna go down to none. One detail matters if you already run Opus in production: older Opus models started at high, so any automation that never sets the effort now runs one level lower than before.
Anthropic's advice is to start at medium and climb one step at a time, and OpenAI's guide also calls medium the right default for most work. I put both labs' advice into one picture you can keep as a reference.

Anthropic also claims that Opus 5.5 at medium beats GPT-6 Astra at max on an office-work benchmark, for about a fifth of the cost per task. Those are Anthropic's own numbers, so I treat them as a claim to check, but they show how much money sits in this one setting.
My advice is to set the effort level yourself on every automation instead of trusting the default. Start at medium, run twenty real examples, then run the same twenty at low and keep low wherever the answers hold up. Only raise the level for the tasks that actually fail, and never leave max on as a standing default. We have not yet measured what each step costs on our own work, so take this as a method to test rather than a result.
The models keep getting smarter at the lower settings, so the real question is shifting from which model to buy to how much thinking each task deserves.
- https://platform.claude.com/docs/en/build-with-claude/effort
- https://www.anthropic.com/claude-opus-5-5
- https://developers.openai.com/api/docs/guides/reasoning
- https://developers.openai.com/api/docs/models/gpt-6-astra
- https://developers.openai.com/api/docs/models/gpt-6-sol
- https://developers.openai.com/api/docs/models/gpt-6-luna
Get the next one by email.
Free every week: what’s new, how to apply it, and the numbers.