Claude Opus 5 shipped this morning, priced the same as the one I trusted, and marketed as a clear upgrade. Here's the three-step check that kept the switch from breaking things silently, and why I hand that check to my AI team instead of running it myself.
Claude Opus 5 shipped this morning. Same price as the one I'd been running. Every early signal said clear upgrade. My first instinct, honestly, was to just switch everything over today, before lunch.
I didn't. The reason I didn't is the whole point of this piece — your model knowledge expires on day one.
Here's the problem with trusting that instinct. Everything I know about how a model behaves, its quirks, its defaults, what breaks it, I learned from models that came before this one. The new model postdates all of that. It cannot be described accurately by my last mental update of “how models work,” because it is the update. It can't even describe itself reliably. Ask a brand-new model to explain its own defaults and you're asking it to be a more current source than it actually is.
So flying in on instinct means flying blind, no matter how confident the instinct feels.
The Three Checks Before You Touch Anything
Before I let anything touch the new model, my own workflow or anyone else's, I run three checks. Not as ceremony. As the only way to see what instinct can't yet.
Research it live. Against the maker's current documentation, not against what I remember about models in general. Memory is exactly the thing that's stale here — the model shipped after my last update to “how models work,” so my update is the thing that needs updating.
Map the blast radius. Separate what I actually control, my prompts, my configuration, my settings, from what the platform now decides for me by default. A new model doesn't just answer differently. It can quietly change who's driving.
Audit for silent breakers. The failures that don't throw an error. The output that looks fine and isn't, right up until someone downstream notices it wasn't.
What the Check Caught
This morning's check found exactly that kind of silent breaker. The new model reasons by default now, and that reasoning shares the same length budget as its final answer. A reply that used to fit comfortably can now get cut off mid-thought, with no error and no warning. Anyone who switched over on instinct would have started getting truncated answers and had no idea why.
It also found a configuration that had worked cleanly on the old model and now hard-fails outright on the new one, a setting that would have needed to change before a single real request could succeed.
And it turned up something nobody was even looking for: a tool already in daily use had been quietly truncating its own answers for weeks, on the old model, the one everyone had trusted the whole time. Nobody had caught it. The only reason it surfaced now is that looking closely at the new model meant finally looking closely, period.
Because the check ran before the switch and not after, the cutover was clean. No broken output. No one chasing a ghost, wondering why a reply just stops. Just a model that worked, and a fix on the model I was leaving behind that nobody knew it needed.
Delegate the Evaluation. Keep the Decision.
Here's the part that isn't really about this one model. My read of a day-old release isn't reliable evidence. It can't be. The model is newer than everything I know about how models behave. That's not a knock on my own judgment. It's just what day one means, for anyone, on any release.
So I don't evaluate on instinct, and I don't skip the evaluation because the upgrade looks obvious on paper. I hand the grounded check, the live research, the blast-radius map, the breaker audit, to my AI team, because that's the part a human can't do well from memory. What I keep is what was always mine: setting the intent, and making the call.
This morning I still switched. I just switched knowing what I was switching into.
— Chris

