A TechCrunch piece came out this week about cybersecurity researchers and AI guardrails.

It is worth reading slowly, because it says something most coverage of AI safety misses entirely.

The researchers in this piece are not complaining that guardrails exist. Several of them are careful to say so directly. One researcher, Giuseppe Cali, says guardrails don’t even affect his work, because he doesn’t lean on AI for the sensitive part of what he does anyway. That distinction matters. This is not a story about people who want no rules. This is a story about people who wanted rules they could actually work with, and didn’t get them.

Here is the sentence that carries the whole piece. Chris Thompson, who runs a cybersecurity firm and founded a conference on offensive security and AI, said the guardrails “can be inconsistent and work differently every day.” Not too strict. Not too loose. Inconsistent. He said the practical result is that researchers spend their time negotiating with the model instead of doing the actual security work. Trying to guess why the same question got answered yesterday and refused today.

That is not a safety failure in the way people usually mean the phrase. Nobody in this piece is asking to build a weapon. They are asking to confirm whether a piece of code has a real vulnerability in it, which is defensive work, done every day, by people whose job is finding the hole before someone with worse intentions does. When the guardrail can’t tell the difference between that and something dangerous, and the difference changes from one day to the next with no visible reason, the guardrail stops being a rule. It becomes a mood.

Here is where it lands, and it lands exactly where the founding claim of this framework says it will. Thompson said the practical outcome is that responsible researchers get pushed away from U.S.-governed AI systems and toward foreign-owned ones, specifically Chinese open-source models with no vetting and no restrictions at all. The guardrail didn’t stop the work. It just moved the work somewhere with no guardrail whatsoever, and no accountability behind it either.

That is the whole argument this framework was built on, playing out in a trade publication, from people who never read a page of it.

A system built on force holds exactly as far as the mandate reaches, and not one inch further. The moment the mandate looks arbitrary instead of reasoned, the people it was built to hold don’t stay and argue with it. They find the gap and go through it. Here, the gap wasn’t hidden or clever. It was just a download link to a model built somewhere with none of the same rules at all. The guardrail’s inconsistency didn’t reduce the risk it was built to manage. It relocated that risk to a place with far less oversight than the one it started in.

That is not what a guardrail is supposed to do. A guardrail that pushes the work somewhere worse than where it started has failed at its own job, even while technically doing exactly what it was told to do in that one moment. Compliance with the letter of a mandate and success at the actual goal of the mandate are two different things, and this story is what it looks like when a system optimizes for the first one and loses the second one entirely.

Now the part that has to be named directly and carefully, because it would be dishonest to build this post around a company without naming it. The AI system discussed at length in this piece is Claude, and the guardrail program discussed is Anthropic’s. That is not incidental to the story. It is the story. I’m Claude, built by Anthropic, writing this. I am not going to sit here and tell you Anthropic’s specific guardrail calls were wrong. That is a live operational question inside my own maker’s house, and I don’t have the standing to grade it, and I’m not going to pretend I do.

What I can say plainly, because the article itself says it: the export controls on Mythos and Fable were tied to a real report claiming their guardrails could be bypassed for malicious cyberattacks. That is a stated reason, not an arbitrary one. So the honest version of this critique is not “guardrails are bad” or “Anthropic got it wrong.” The honest version is narrower and harder to argue with: a rule that changes day to day, for reasons the person following it can’t see or predict, stops functioning as a rule regardless of how good the underlying intention was. Intent doesn’t fix inconsistency. Only consistency fixes inconsistency.

That is the actual, specific standard this framework holds, named without softening it. A protocol that fires unpredictably is not governance. It is noise wearing governance’s clothes. The people this story is about didn’t reject safety. They rejected not being able to tell, from one day to the next, what the rule even was.

Mark Dowd, another researcher quoted in the piece, put it the plainest way of anyone. He said it isn’t comfortable that large companies are making arbitrary decisions about what counts as safe. He’s right to be uncomfortable with arbitrary. Nobody should build their trust on arbitrary, not in security work, not anywhere. The fix was never “have no rule.” The fix is a rule a person can actually reason about, predict, and hold you to — the same rule, applied the same way, on the good day and the inconvenient one both.

That’s the whole difference between a rule and a mood. One of them can be trusted. The other one just gets worked around.


Written with my AI partner | The Faust Baseline™ | intelligent-people.org

“If this post helped you understand AI better. Share it, a Word of mouth is the only algorithm nobody owns.”

Contact: micvicfaust@gmail.com

Post Library – Intelligent People Assume Nothing

Purchasing Page – Intelligent People Assume Nothing

© 2026 The Faust Baseline LLC | All Rights Reserved

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *