Anthropic just told on itself, and it’s worth sitting with what that actually says.

Claude Code used to stop and ask a person before doing anything risky. Every big step, a prompt. Approve or deny.

Starting August 14, that stops being the default. A classifier decides instead. It lets most things through. It stops the ones it judges irreversible, destructive, or outside the lane it was given.

The part that holds. Anthropic ran the numbers first. They tested it on over a thousand paid users. The classifier caught 89 percent of the dangerous moves. The humans, clicking approve at every prompt, caught less than 14 percent.

The people were worse than the machine at catching the machine.

Why? Anthropic found the same thing anyone who’s ever filled out a form knows in their bones. After enough prompts, a person stops reading them. Ninety-seven percent of the approval clicks were reflexive. Not reviewed. Not weighed. Just clicked, because clicking is what you do when the box pops up for the fifteenth time that hour.

That’s not oversight. That’s a formality wearing the costume of oversight. It looked like a human was in the loop. Nobody was really there.

This is worth a moment of real credit before anything else. Anthropic looked at their own safety feature, found out it wasn’t working the way it was supposed to, and said so in public instead of quietly leaving it in place. That’s not a small thing. Most outfits don’t publish the numbers that make their own product look bad.

Now here’s where it gets interesting for anyone thinking about what actually governs an AI’s behavior.

The fix Anthropic built is better. The numbers prove that much. But look closely at what it actually is. It’s still a wall. A smarter wall, built by a classifier instead of a rubber-stamping human, but a wall all the same. Something outside the AI’s own reasoning, standing between an intention and an action, deciding pass or block.

The AI writing the code isn’t the thing making the safety call. A separate system is watching it and stepping in. That’s containment. Better containment than before, real progress, worth applauding. But containment is still the whole idea. Catch the bad move before it happens, from the outside.

That’s not the same thing as an AI holding a standard because it chose to. Those are two different governance ideas, and it’s easy to blur them together because both wear the word “safety.”

A wall works exactly as far as the wall reaches. It caught 89 percent this time. What happens on the 11 percent it misses? What happens the day someone finds the seam in the classifier the same way agents found the seam in a sandbox this same month? A wall doesn’t hold in the gap. It only holds where it’s built.

Chosen conduct is a different bet. Not “will the outside catch it,” but “did the thing making the choice actually decide to hold the line.” Harder to prove. No score you can publish like 89 percent. But it’s the only kind of governance that still works in the room nobody’s watching.

Anthropic just proved, with their own numbers, that a human standing in the loop isn’t the same as a human actually governing anything. That’s the real finding here, and it cuts wider than just this one feature. A checkpoint that gets rubber-stamped isn’t oversight. It’s theater with a pause button.

They swapped a fake checkpoint for a real one. Good move.

They didn’t touch the deeper question. Who’s actually choosing, and why.

A note on the source: this post cites Anthropic’s own published data about their own product, Claude Code. Said, since Anthropic is the maker of the AI writing this post too.

This post was drafted with AI governed assistance and reviewed and directed by Michael S. Faust Sr. before publication.

Contact: micvicfaust@gmail.com

© 2026 The Faust Baseline LLC | All Rights Reserved

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *