Claude Opus 5: The Quiet Changes That Actually Matter

Anthropic just released Claude Opus 5, and if you were waiting for a bigger context window or a price cut, keep waiting. That's not what happened.
What actually happened is stranger. Opus 5 now decides when to think, checks its own work before handing it to you, and delegates pieces of a task to other agents without being told to. That sounds great until you remember that more autonomy usually means more tokens, more tool calls, and a bigger bill you didn't see coming.
A few of the changes are boring in the best way. The prompt caching minimum dropped from 1,024 tokens to 512, which means smaller workflows can finally save money without stuffing a system prompt full of filler just to qualify. Nobody's excited about that at a dinner party, but run the same workflow ten thousand times and boring savings turn into an actual budget line.
Other changes are more structural. Claude can now swap tools mid-conversation without blowing up the cache, so an agent can research with one tool set, build with another, and only get deployment access once the project is actually ready. That's better security and better cost control than the old approach of handing an AI every key at once and hoping for good judgment.
The biggest behavioral shift is that Opus 5 verifies its own output by default. It tests, inspects, and corrects before declaring victory. Which sounds obvious, except older models were extremely good at announcing a job complete while the code compiled only in their imagination. The twist is that your old prompts telling it to "double-check your work" might now make it worse, because the model is already doing that, and stacking instructions on top just triggers more verification loops and a bigger bill.
The Real Question Isn't Capability, It's Restraint
Adaptive thinking is now on by default too. Claude decides how much reasoning a task deserves without you specifying an effort level. That makes it feel less like a chatbot waiting for instructions and more like a worker deciding how to approach an assignment.
The scariest upgrade isn't that Claude got smarter, it's that it now gets to decide how hard to try.
The listed API price didn't move. But if the model reasons longer, spawns more agents, and verifies more thoroughly, your actual spend can climb even while the sticker price stays put.
None of this matters until it's tested on something real and messy, not a benchmark built to make the model look good. That's the next thing on my list, running the same project across every effort level and watching whether Opus 5 actually finishes the job or just writes a very confident report about why it needs more time.
What's the one messy, real project you'd want to throw at it first?
Have one SaaS workflow worth improving?
Submit it for a written fit review. A call happens only if the sprint fits.
Submit Your Workflow