A security researcher just ran the biggest test yet of a question the whole AI governance conversation keeps dancing around.
Can you let AI loose to hunt for vulnerabilities on its own, or does it need a human holding the leash?
James Kettle runs research at PortSwigger, the company behind Burp Suite. He built a system called HTTP Terminator. He fed it 138 technical documents on how web servers are supposed to talk to each other. The AI broke those down into 15,000 fragments and used them to dream up 30,000 possible ways to break that conversation.
Then it tested every one of them. Live. On real websites that had agreed to be tested.
Out of 30,000, about 700 came back vulnerable. Banks. Government systems. Security products. An airport.
Here is the part that matters for the Baseline.
Kettle did not build an AI and walk away. He built the loop, then stayed inside it. Deterministic code — plain, rule-following, no-judgment-calls code — decided what the AI was allowed to act on by itself. A human decided which findings got pushed further. A human validated the real discoveries before anyone called them real.
Kettle said it himself. This inverts the story everyone has been telling about AI research. The value was not the AI working alone. The value was a human amplified by it.
That is not a new idea to anyone who has read this far into the Baseline.
AGP-1, the Agentic Governance Protocol, was built on that exact premise. An AI system acting with any independence needs boundaries a human sets and a human checks. Not because the AI cannot be trusted with everything, but because oversight is the thing that turns raw capability into something a person can stand behind.
Kettle did not read AGP-1. He did not need to. He ran the experiment, hit the wall on full autonomy, and built the human checkpoint back in — because the work demanded it, not because a framework told him to. That is the strongest kind of proof there is. The pattern shows up on its own, from someone with no reason to go looking for it.
One correction before this goes further.
An earlier version of this research note included a claim that Anthropic’s Mythos model had found 231 Microsoft vulnerabilities faster than patches could follow. I went back and checked that claim against the source reporting. I could not find it anywhere — not in the PortSwigger research, not in any of the coverage around it. It does not check out. It is cut from this post, and it should not have been in the draft that reached me.
That is worth naming plainly rather than quietly fixing. A governance framework that catches its own bad claims and says so, in public, is doing exactly what it is supposed to do. One that quietly edits and moves on is not.
So where does this leave things.
The security world just got a real-world case study showing that AI research without a human check produces noise, and AI research with one produces a new vulnerability class nobody had named before — Shared-Parser Confusion, something neither the AI nor Kettle alone would likely have found.
Full autonomy did not win. Full human control did not win either. The mix won.
That is the whole argument this Baseline has been making since before most people were asking the question. Not cage the AI. Not turn it loose. Build the conduct standard, and let a human and a system that has agreed to follow it work the problem together.
One hacker in a bug bounty program just handed the field free evidence for it.
This post was drafted with AI governed assistance and reviewed and directed by Michael S. Faust Sr. before publication.
Contact: micvicfaust@gmail.com
© 2026 The Faust Baseline LLC | All Rights Reserved






