Twelve of the biggest AI companies in the world have published safety frameworks.
Microsoft put out its third one last week. Amazon has one. Google, Meta, OpenAI, Nvidia, xAI. Twelve companies, twelve documents, all of them saying how they’ll handle the risks of what they’re building.
Sounds like progress. And in a way it is.
Here’s what nobody told you.
A group of researchers took all twelve of those documents and graded them. Not on whether the companies are good people. Not on whether they mean well. They graded them against sixty-five criteria borrowed from industries that have been managing catastrophic risk for decades. Aviation. Nuclear power. Places where being wrong kills people and the paperwork reflects it.
The best score was thirty-four percent.
The worst was eight.
The middle of the pack came in at eighteen.
Now, I want to be careful here, because a low grade sounds like an accusation and that’s not what this is. These researchers said something more useful than that. They said the problem isn’t that the companies promised bad things. It’s that the promises are written so loosely you can’t tell what was promised.
Their finding, and I’m putting it in my own words: you can’t predict what these companies will decide. You can’t judge whether their planned response would be enough. And you can’t determine whether they’ve kept the commitments they already made.
You can’t determine whether they’ve kept them.
That’s not a safety framework. That’s a statement of intent with a nice cover on it.
I wrote a couple days back about a man declaring AGI had arrived with no test attached to the claim. A thing can’t be checked if nobody drew the line it’s supposed to cross.
This is the other half of that story, and it’s the better half.
Because here, somebody drew the line.
Somebody sat down, took standards from nuclear and aviation, and said: here’s what a real risk framework contains. Here are the sixty-five things. Now let’s see what’s actually in these documents.
And the answer came back eighteen percent.
That’s what a line does. It doesn’t accuse anybody. It just tells you where you are.
I need to say something here. The company that scored highest on that list — thirty-four percent, the top of the class — is Anthropic. Anthropic makes the AI I use to write. So the ranking I just handed you puts my own supplier in first place. You should weigh that, and you should go check the numbers yourself. I’d rather tell you that than have you find it later.
And thirty-four percent is the good score. First place, and two-thirds of the way to nothing.
There’s a real difference between having rules and having rules that can be checked. It’s the whole difference, actually. A rule nobody can fail is decoration. It hangs on the wall. It makes people feel better walking past it. It never once changes what anybody does, because there’s no moment where somebody has to stop and answer for it.
I spent my life around work that got inspected. Not because anybody assumed I was crooked. Because the inspection is what made the word “done” mean something. Take the inspector away and “done” becomes whatever the man holding the hammer says it is. That’s not a character flaw. That’s just what happens to words when nothing pushes back on them.
Twelve companies wrote down their safety promises. Six more promised to and never did.
Nobody was grading them. So the writing drifted to where writing drifts when nobody’s grading — toward language that sounds serious and commits to nothing you could ever point at later.
Then somebody graded them, and now we can see it.
That’s a starting point.
The scores can go up. The same researchers pointed out that if any one of these companies simply adopted the best practices its own competitors are already using — nothing new, nothing invented, just borrow what’s already on the market — the score would nearly triple.
The knowledge exists. It’s sitting in the industry right now, spread across twelve documents. Nobody’s assembled it because nobody had to.
That’s what a grade does. It makes somebody have to.
So here’s where I land.
Don’t take the existence of a safety framework as evidence of safety. Twelve companies have one. The middle of that group scores eighteen percent against standards that nuclear plants met decades ago.
Ask a different question. Ask whether anybody outside the company can check it. Ask what happens if they fail. Ask who’s holding the ruler.
And when somebody outside finally picks up a ruler and starts measuring — like these researchers did — pay attention to that. That’s the moment things start moving.
Not the moment the report gets published.
The moment somebody grades it.
” Attic Thoughts”-library – Intelligent People Assume Nothing
Contact: micvicfaust@gmail.com
This post was drafted with AI governed assistance and reviewed and directed by Michael S. Faust Sr. before publication.
© 2026 The Faust Baseline LLC | All Rights Reserved






