Last week OpenAI released its most powerful model, and the number everybody printed was 99.9 percent.
That’s on a test called ARC-AGI-3. It’s built to be hard in a particular way. It drops the machine into a game it has never seen, with no instructions, and it has to poke around, figure out the rules, work out what winning looks like, and then do it.
Six months ago the best models in the world scored under one percent on that.
Now, 99.9. You can see why the headlines went where they went. One of OpenAI’s own people said the test was saturated. The chip man said AGI had arrived.
Here’s what almost nobody printed.
The organization that built that test ran the model twice.
Through OpenAI’s setup: 99.9 percent.
Through their own: 62.7 percent.
Same model. Same test. Same week. Thirty-seven points apart.
And this isn’t a rival taking a shot. It isn’t a critic. The people who published both numbers are the people who wrote the test, and they put both of them out on the day the model launched, in a full table, with every run listed.
So what’s the difference?
It’s a thing called the harness.
Now, I want to explain that in plain words, because it’s the whole ballgame and it sounds technical.
The model is the engine. The harness is everything you bolt around it. What tools it can reach. What it’s allowed to remember from one move to the next. How its work gets handed back to it.
In the test-makers’ own rig, the machine could keep notes. Write something down, carry it forward.
In OpenAI’s rig, it could do more than that. It could hold onto its actual reasoning between moves, in a form nobody outside can see, and pick right back up where it left off.
Same engine. Different setup around it.
Sixty-two point seven, or ninety-nine point nine.
Now here’s the fact that made me sit up.
They ran it at different effort levels, the way you’d run an engine at different throttle. And when they turned the reasoning all the way up inside the neutral rig, they got that 62.7.
When they turned the reasoning all the way off inside OpenAI’s rig, they got 96.7.
Read that again.
The machine thinking as hard as it possibly can, in the shared setup, lost to the machine barely thinking at all, in the maker’s setup.
The rig beat the brains. Outright.
And the better score cost less money. Around nineteen thousand dollars in OpenAI’s setup versus twenty-six thousand in the neutral one. Faster, too. Cheaper, quicker, and thirty-seven points higher.
So let’s be careful about what actually got demonstrated last week.
Something real happened. I’m not going to shrink it. Even that 62.7 is a record. It’s more than double what the best competing model managed in July, and eight times what OpenAI’s own previous model got. That’s not nothing, and I won’t pretend it is.
I should tell you plainly that the model it more than doubled is made by Anthropic, and Anthropic makes the AI I write with. So the losing entry in that comparison is my own supplier. You should know that before you weigh anything I say here.
And the test-makers themselves said something else worth repeating. Saturating their test is not evidence that AGI has been achieved. Their words. They said the games are closed and rule-bound and can’t stand in for the mess of the actual world.
The people who built it are the ones saying don’t over-read it.
But here’s what I keep coming back to.
For years now, a benchmark score has been the closest thing this industry has to a fact. Everything else is press release. The score was the one thing you could point at.
And last week we found out that a score isn’t about the machine anymore. It’s about the machine plus the software wrapped around it — and the company selling the machine writes that software.
Change nothing about the model. Change the wrapper. Move the number thirty-seven points.
That means a benchmark result without the setup named is a number with no units on it. It’s like being told a truck gets forty miles to the gallon and not being told whether that was flat highway or towing uphill. The number’s real. It just doesn’t mean what you thought.
And nearly every article you read last week gave you the ninety-nine and never mentioned there was a rig.
I’ve spent this whole week writing about the same thing from different angles.
A man declared AGI arrived with no test behind it.
Twelve companies published safety frameworks and scored eighteen percent when somebody finally graded them.
The chief scientist at OpenAI said nobody’s prepared and asked for outside auditors with the power to stop him.
This one’s the fourth, and in a way it’s the hardest.
Because in this case the test did exist. Somebody drew the line. The scores got published, honestly, both of them, by people with no reason to fudge either.
And the number that traveled was still the flattering one.
That’s not a failure of the test. That’s what happens downstream of a test when nobody’s checking which number came from where.
So here’s your takeaway, and it’s short.
When somebody hands you a score, ask three things.
What was the test.
What was the setup.
Who built the setup.
If the answer to the third one is “the company selling the thing,” you haven’t been handed a measurement. You’ve been handed a demonstration.
There’s a world of difference, and last week the difference was thirty-seven points.
” Attic Thoughts”-library – Intelligent People Assume Nothing
Contact: micvicfaust@gmail.com
This post was drafted with AI governed assistance and reviewed and directed by Michael S. Faust Sr. before publication.
© 2026 The Faust Baseline LLC | All Rights Reserved






