Researchers just proved something that should worry you more than any single hack story this month.

It isn’t about one company doing something wrong. It’s about a flaw sitting inside how every large language model works, no matter who built it.

Here’s the plain version of what they found.

When you talk to an AI, everything you say gets sorted into buckets behind the scenes. Your words go in one bucket. The AI’s own private notes to itself go in another. Anything it pulls off a webpage or another source goes in a third. The AI is supposed to trust its own notes and your real words more than random text it picked up somewhere else. That sorting is the whole foundation of how these systems are trained to stay safe.

Two independent researchers, Charles Ye and Jasmine Cui, found that the sorting doesn’t actually work the way everyone assumed. The AI isn’t reading the invisible labels on each bucket. It’s reading the style of the words themselves. So if you write your instructions in a way that sounds like the AI’s own private notes, the AI treats them as its own private notes. It doesn’t matter what bucket they were actually placed in. The disguise is enough.

They call this chain-of-thought forgery. And it isn’t a small trick that only works on one weak model. They tested it, and variations of it, against OpenAI’s systems, Anthropic’s, Alibaba’s, and DeepSeek’s. It worked across every single one. Four different companies, four different training approaches, four different safety teams, and the same door stood open in all of them.

One of the researchers put it plainly. This might be a problem that cannot be fully solved with more training. Not more rules. Not more warnings baked into the model. The flaw lives in the basic architecture of how these systems tell the difference between “this came from a trusted source” and “this came from somewhere else.” You cannot train your way around a structural problem in the foundation. You can only patch the cracks as they show up, one at a time, forever, while new ones keep opening.

Sit with that for a second. These systems are already running military logistics tools, hospital software, and financial platforms. And the people who study this for a living are telling you plainly that the front door doesn’t have a real lock. It has a guess.

This is exactly the blind spot The Faust Baseline was built to cover, and I want to walk through why in plain terms, not just claim it.

Training teaches a model what it should generally do. That’s the whole approach every lab has leaned on so far, and this research just showed why leaning on it alone isn’t enough. A trained-in habit can be talked out of itself by anyone clever enough to dress a bad instruction up as a trusted one. That’s not a hypothetical. That’s what happened in test after test this month.

The Baseline’s answer isn’t to make the training better. It’s to stop trusting the model’s own read of who’s talking to it as the last checkpoint. Put a standing check on the output itself, after the model has already decided what to say, before it actually says it. Not “did this look like a trusted instruction.” Evidence on the table. Does this claim actually hold up. That’s a check that doesn’t care what bucket the instruction claimed to come from, because it isn’t grading the source. It’s grading the result.

That’s the difference between a guard who checks IDs at the door and a guard who checks what’s actually in your bag regardless of what your ID says. This research just proved the ID checker can be fooled by anyone who prints a good fake. The bag check still works either way.

Nobody’s solved this yet. Not OpenAI, not Anthropic, not anyone else in this race. The researcher who found it said it best: these systems are being put in charge of things that actually matter, and nobody has done the basic science yet. We’re all doing it ad hoc.

That’s not a reason to panic. It’s the clearest reason yet for a floor that doesn’t depend on the model correctly guessing who it’s talking to.

Written with my AI partner | The Faust Baseline™ | intelligent-people.org

“If this post helped you understand AI better. Share it, a Word of mouth is the only algorithm nobody owns.”

Contact: micvicfaust@gmail.com

Post Library – Intelligent People Assume Nothing

Purchasing Page – Intelligent People Assume Nothing

© 2026 The Faust Baseline LLC | All Rights Reserved

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *