Here’s a story making the rounds in security circles this week. It’s short. It’s simple. And it should scare anybody betting the farm on AI running unsupervised.
A security research group called 1Password Off-by-1 Labs decided to test something. They took two of the leading AI models and gave them a job. Find security holes in code. Fix them. Write the patch yourself.
They didn’t run one test. They ran over six thousand of them. 6,080 patches, generated across a wide range of vulnerability types. Real bugs. Real fixes. Real grading afterward.
Only 26 out of every 100 patches were clean. Fully fixed. No holes left behind.
That means 74 out of 100 were not clean. Three out of every four patches an AI wrote to fix a security hole failed to actually close it properly.
In roughly half of every case tested, the patch left at least one way in. The vulnerability the AI was supposed to seal shut stayed open. A hacker could still walk through the same door, even after the “fix” went in.
The researchers didn’t stop at giving the AI good information. They also tested what happens when the AI starts with bad information. Wrong guidance about what the bug actually was.
When that happened, the success rate didn’t just drop. It collapsed. Down to about 15 out of 100.
The AI didn’t catch the bad guidance and correct itself. It didn’t pause and say something’s wrong here. It took the wrong starting point and built a wrong answer on top of it, with the same confidence as if it had been right all along.
Not the failure rate by itself. The confidence sitting on top of the failure rate.
An AI that gets it wrong and knows it got it wrong is a tool you can work with. An AI that gets it wrong and hands you the wrong answer wrapped in the same tone as the right answer, that’s a different animal entirely. That’s the one that gets a business in trouble.
The researchers were blunt about their conclusion. Letting AI patch security holes on its own, without a human checking the work, is worse than not automating at all. The harm from broken or missed patches outweighs whatever speed you gained by skipping the human step.
The experts who ran this study are not anti-AI. They ran six thousand tests because they wanted to know if the tool was ready to work alone. The data came back and told them no.
This is the exact ground The Faust Baseline has been standing on since day one. Protocols are chosen conduct, not self-enforcing architecture. An AI does not govern itself just because you asked it nicely to be careful. Somebody has to actually check the door before you tell the customer it’s locked.
The Baseline calls this out directly in two places. BLP-2 is the Boundary and Reasoning Integrity Protocol. It says an AI has to name the wall in front of it, not pretend the wall isn’t there. AGP-1, the Agentic Provision Standard, says the same thing from a different angle. When an AI is handed the wheel to actually do something, not just talk about it, somebody has to build the guardrail in from the outside. The AI checking its own work is not a guardrail. It’s the same student grading his own test.
This is not a small corner of the tech world either. Compliance teams are already circling this data. Under things like the EU’s Cyber Resilience Act, a security patch that leaves a known hole open is not a technical footnote. It’s a liability. It’s an audit finding. It’s the kind of thing that shows up in a courtroom later, not just an engineering post-mortem.
So what do you do with a number like 26 percent?
You don’t throw the tool out. That’s not the lesson here. The lesson is smaller and harder to argue with. You don’t let the machine be the last set of eyes on its own work. Not on something that can let a stranger into your system. Not on anything that matters.
A human has to be the one who says, this is actually fixed, before it goes out the door. That’s not a limitation on the technology. That’s the whole point of having a human in the room at all.
The tools are getting faster every year. Faster is not the same as trustworthy. This study is one more piece of proof, sitting right out in the open, that the two are not the same thing.
Challenge line. If I’m wrong here, tell me where. Maybe 26 percent is actually a good number for a first-generation tool doing something this hard, and the real story is how fast that number climbs from here, not how low it sits today. That’s a fair argument. I’d want to see the next study before I let go of it.
This post was drafted with AI assistance and reviewed and directed by Michael S. Faust Sr. before publication.
Contact: micvicfaust@gmail.com
© 2026 The Faust Baseline LLC | All Rights Reserved






