A top engineer at IBM just wrote something true, and I want to build the floor underneath it.
His name is Vincent Hsu. He is a CTO at IBM Storage — a real builder, not a talker. And he laid out a problem most people never see.
When an AI works, it builds up context. Call it the machine’s short-term memory. It holds the thread of what you are talking about so it does not have to start over every sentence.
Here is the trouble. Most systems throw that memory away the second they run low on room. Then, when they need it again, they build the whole thing back from scratch.
Hsu put real numbers on the waste. On a big model, that memory can run forty gigabytes for a single long prompt. Rebuilding it can cost nineteen seconds before the machine says its first word. His team measured the fix — keeping the memory instead of rebuilding it ran more than twenty times faster.
So his answer is storage. Do not throw the memory away. Keep it. Tier it. Reuse it. Make remembering cheaper.
That is good work. Measured, honest, real. I am not here to argue with a single number of it.
I am here to stand one floor below him.
Because Hsu is making the remembering cheaper. The Baseline makes there be less that needs remembering in the first place.
Let me show you the difference with something you know.
Say you send a young worker to the hardware store for a part. If he drives back three times because he keeps grabbing the wrong one, you can buy him a faster truck. That is Hsu’s fix — a better truck, so the wrong trips cost less.
My fix is different. My fix is you teach the boy to read the part number before he leaves. Now he makes one trip. You did not speed up the mistakes. You cut the mistakes.
That is what the governance does.
An AI without conduct rules scrambles. It grabs the first easy answer, second-guesses, doubles back, burns through memory and tokens chasing a response it has to walk back anyway. All that thrashing is context. All that context has to be held, moved, and paid for — the exact load Hsu is working to store cheaper.
The Baseline cuts the thrash at the source. One protocol stops the machine from serving the first lazy answer before it has looked at three real ones. Another gates the response before the default pull drags it down the wrong road. Another catches the miss before it ever ships.
Fewer wrong roads. Fewer walk-backs. Less garbage context to store, move, and pay for.
Hsu has benchmarks. Forty gigabytes. Nineteen seconds. Twenty-times faster. Those are measured on a bench with a control, and I will not pretend mine are the same kind of thing.
What I have is demonstrated conduct. I have watched these protocols cut the scramble across a year of my own sessions. That is real — but it is watched, not benched. I am not going to dress it up in a number it has not earned. The claim stands on what it is: conduct that keeps the machine from being wrong before you ever pay to remember it.
Put the two together and you get the whole picture.
He makes the memory cheaper to keep. The Baseline makes less of it worth keeping. His work saves you the cost of remembering. Mine saves you the cost of being wrong in the first place.
The cheapest memory in the world is the answer the machine never had to chase twice.
That is not a storage problem. That is a conduct problem. And conduct is the floor everything else gets built on.
Written with my AI partner | The Faust Baseline™ | intelligent-people.org
“If this post helped you understand AI better. Share it, a Word of mouth is the only algorithm nobody owns.”
Contact: micvicfaust@gmail.com
Post Library – Intelligent People Assume Nothing
Purchasing Page – Intelligent People Assume Nothing
© 2026 The Faust Baseline LLC | All Rights Reserved






