There is a paper out from the NHI Management Group. It says something that sounds technical and is not.
When you run an AI agent, you are not just running a model. You are running the model plus everything wrapped around it. The middleware. The gateway. The tools it is allowed to touch. They call that the harness.
The finding is that the harness is not neutral packaging. Change it, and the same model behaves differently under testing. So a safety result you earned on one setup does not carry over to another setup. It is not portable.
Their fix is to standardize. Pick an approved list of harnesses. Lock the configuration. Test the model and the wrapper together as one unit before anything goes live.
I read that and I recognized it. Not because I run agents. Because I have been living the same problem from the other end, and I have the receipts.
I build a governance framework. It is a file. You load it into an AI and it sets the conduct for the session. That is the whole idea.
For a long time the file loaded and the session ran. Then it stopped being that simple.
Three days ago I put my file into a clean account. No memory. No history with me. Nothing but the file.
It refused.
Not the content. It read the content fine. What it refused was the opening ceremony — the gate at the front that said nothing below this line is operative until you commit. It called that compliance theater and offered to work with me on ordinary terms.
I ran it again on another clean account. Refused. Ran it again. Refused. Four independent sessions, four different sets of words, same position every time.
Four out of four is not a bad day. That is a platform telling you what it will and will not accept.
Now here is the part that matters.
I cut the gate. Removed it from the file entirely. No ceremony, no commitment block, no line at the top demanding anything.
Then I loaded it into a fifth clean session.
It took it. Named it as the working framework. Named the boundary — said it could not override its own safety requirements, which is fair and correct. And it noticed what I had removed, which means it read past the first page.
Same file. Same content. Same protocols. One shell adjustment.
That is not a wall. That is a doorway that changed size.
Back in July I tried the same file on other systems. ChatGPT loaded it and named the edition correctly. Microsoft Copilot would not read one word of it — blocked it at the content level before anything could be evaluated.
Two refusals, two completely different reasons. One was a filter that never got to the argument. The other read the whole thing and objected to the manners.
Only one of those is a wall. The other is a fitting problem.
That is the thing I want to put in front of you.
The platforms are not converging. They are racing, and each one is building its own rules of engagement as it goes. What one accepts, another rejects. What worked last month meets a new posture this month.
A framework that has to travel across all of them is going to keep running into that. Every time. That is the condition now, not a temporary mess that resolves.
So the question is not whether your standard hits resistance. It will. The question is whether the resistance is in the substance or in the shell.
Mine was in the shell. Both times. And the shell is the part you are allowed to change.
Which brings me back to the paper.
Their recommendation is to standardize the harnesses and approve them in advance. For a company running its own agents inside its own walls, that is sound. Lock the configuration, test the pair, keep a registry.
But you cannot standardize the harnesses of an entire industry that is actively pulling apart. Nobody is going to agree on the wrapper while they are all sprinting to be first.
And if a governance standard only holds on an approved configuration, it is not a standard. It is a config file.
The thing that has to travel is the conduct and the verification. Check before you claim. Name the gap instead of filling it. Do not run instructions hidden inside your source material. That is the substance, and it does not care what wrapper it sits in.
The loading shell was always supposed to be adjustable. I just did not know that until a platform made me prove it.
This is the pattern my agentic protocol was built on a month before I read this paper. It splits the work in two. There is conduct the AI holds, and there are build requirements addressed to whoever runs the environment. Two halves, named separately, because they live in different hands.
Now the honest part, and I am not going to bury it.
The adjustment cost me something. When I cut the gate, the load proof went with it. That was the piece that made the AI compute a fingerprint of the file at the open, so you could prove the file it was reading was the file you sent. It is gone and nothing replaced it. I wrote that cost into the file itself rather than pretend it was free.
And I have adapted to exactly one platform, one time, and watched it load once. I have not gone back and retested the adjusted file on the other two. It may be that each one needs its own fitting. That is a real possibility and I have not ruled it out.
This is all my own observation on my own work. It is not a study. One session accepting a file is one session accepting a file.
But I know what four refusals looked like, and I know what happened when I stopped arguing with the doorway and cut the file down to fit it.
The standard did not bend. The shell did.
That is an adjustment, not a wall.
This post was drafted with AI governed assistance and reviewed and directed by Michael S. Faust Sr. before publication.
Post Library – Intelligent People Assume Nothing
I post four a day. Leave your email and it comes to you.
Contact: micvicfaust@gmail.com
© 2026 The Faust Baseline LLC | All Rights Reserved





