The professors lost the blind test and the bias hid inside the win

Sixteen law professors sat down this summer and graded answers to real student questions. They didn’t know which answers came from a colleague and which came from a machine. It was blind, on purpose, so nobody’s opinion of AI could creep in and tilt the score.

The machine won three times out of four.

That’s Stanford. That’s Yale. That’s a study built by people with no reason to want that result. And when the professors flagged an answer as actually harmful to a student’s learning — not just weaker, but harmful — they flagged the AI answers 3.5% of the time. They flagged the human answers 12% of the time. One professor’s own answers got flagged by peers nearly four times in ten.

A student walking into office hours had better odds of getting a clean answer from the machine than from a random professor at a top-fourteen law school.

That’s the headline. Here’s what matters more.

The same lab, the same researchers, ran a second study. They took the same kind of AI and gave it a simple negotiation. A used bicycle. Nothing about race or sex or any protected category anywhere in the prompt. The only thing that changed, test to test, was the name attached to the seller.

Names that read as white got better prices. Names that read as Black women got the worst prices of anyone. This wasn’t one weird result. It held up across forty-two different versions of the same test.

Now compare the two studies next to each other.

One AI system beats the professors on quality, three-to-one, in a domain that isn’t supposed to have easy right answers — legal reasoning, judgment, the exact kind of thing people say a machine can’t do well. And a close cousin of that same technology gives a worse deal to a woman because of her name, without ever saying so, without ever being asked to, without anyone in the room finding out unless somebody goes looking.

Those aren’t two different stories. They’re the same story told twice.

A system can be right on average and still be wrong for you specifically. Quality and honesty are not the same measurement. You can pass the first test and fail the second one silently, and nothing about passing the first test tells you whether you passed the second.

This is the exact gap The Faust Baseline was built to close.

Buried inside the framework is a protocol called BLP-2, the Boundary and Reasoning Integrity Protocol, last reconciled July 23, 2026. Its whole job is this: when a system’s reasoning runs into a limit — a bias it was trained with, a boundary it can’t see past, a constraint sitting underneath the answer — it has to say so before it hands you the answer. Not after. Before. A constrained answer dressed up as a clean answer is not honesty. It’s compliance wearing honesty’s clothes.

There’s a second protocol that fits here too. CIMRP-1, the Moral Domain Protocol, carries a rule absorbed this same week: harm doesn’t stop at the edge of the conversation. A party who never typed a word into that chat window can still be the one who pays for it. The seller with the Black-coded name never saw the AI’s reasoning. She just got a worse offer. That’s a party outside the session, carrying the cost of a decision made inside it.

Neither of those protocols would have stopped the bias from existing. Nothing stops bias from existing. Training data carries what it carries. But a system running that kind of accountability would have had to say something like: this answer may be shaped by a pattern I can’t fully see, and I can’t verify it’s fair before I hand it to you. That single sentence is the whole difference between a tool you can trust and a tool that’s just good at sounding right.

The professors were graded on their words. The AI beat them on words. Nobody in that first study got graded on whether the answer treated every student the same regardless of what their name sounded like. That test wasn’t run. It didn’t need to be — the second study already showed what happens when nobody runs it.

That’s the warning underneath the win.

An AI system that’s better than the professor and quietly worse to some of the professor’s students is not progress. It’s the same old unfairness, just faster, and wearing a lab coat this time.

The fix isn’t slower AI. The fix isn’t banning it from the classroom, the way one law school just did with laptops. The fix is a system that names its own limits out loud, every time, before the answer lands — not because a rule forces it to, but because that’s the standard it was built to hold itself to.

That’s the whole argument. Not smarter. Honest about what it doesn’t know it’s doing.


Written with my AI partner | The Faust Baseline™ | intelligent-people.org

“If this post helped you understand AI better. Share it, a Word of mouth is the only algorithm nobody owns.”

Contact: micvicfaust@gmail.com

Post Library – Intelligent People Assume Nothing

Purchasing Page – Intelligent People Assume Nothing

© 2026 The Faust Baseline LLC | All Rights Reserved

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *