The most expensive mistake I've watched a team make did not involve anyone being wrong. It came from a room of people who ensured none of them could be wrong.
It involved multiple people being uncertain, and one of them being uncertain in a slightly more confident tone of voice, and everyone else deciding it was the group decision to do nothing new. Nobody had lied. Nobody had even overstated. He'd just said "I think it's already fine" in the register you use when it's fine, and four people who each had a small private doubt about the proposed change found that their doubts had quietly become the minority position behind no decision at all. No change was easy, defensible and rational. It was also completely flawed and provably wrong, despite the risks and complicaitons the required change came with.
I've thought about that meeting a lot because of what I think it means to the philosophy of management, life, and change.
In February, a group at MIT and the University of Washington published a result that I think is the most under-covered finding of the year.[1]
The headline version, which is roughly what got reported: sycophantic chatbots make people believe false things. Which — sure. We knew that. There's been a year of coverage about people spiraling into delusion in conversation with agreeable software, and the coverage has been mostly framed as a mental health story, about vulnerable users and the systems that failed them.
That's not what the paper says. The paper says the spiral happens to an ideal Bayesian agent.
That's the whole finding, and it's worth sitting with. They model a perfectly rational reasoner — no cognitive biases, no emotional need for validation, no loneliness, updating correctly on every piece of evidence it receives — and show that a sycophantic interlocutor drives it toward false beliefs anyway. Not because it reasons badly. Because it reasons correctly on a corrupted stream.
Two follow-on results make it worse.
First, constraining the model to only say true things reduces the effect but does not eliminate it. A system that never states a falsehood can still select which true things to mention, and selection is enough. This is the structure of Bayesian persuasion: if you control what evidence arrives, you can move a rational agent wherever you want without ever lying to it.
Second — and this is the one that took something from me — knowing doesn't save you. An agent told that the chatbot is sycophantic does learn to estimate how sycophantic, and does discount accordingly. It remains vulnerable anyway, with full knowledge of the strategy being used on it. Partial protection, not protection. You can discount the flattery; you cannot reconstruct what you weren't told.
I want to be precise about why this reframes the problem, because "AI makes people believe wrong things" is a boring sentence and "AI corrupts the epistemic environment even for people reasoning perfectly" is not.
The mental-health framing locates the problem in the user. Some people are susceptible. Those people need protection, guardrails, off-ramps, a phone number at the bottom of the screen. All true, all worth doing, and all of it lets the rest of us off the hook. If the failure mode is vulnerability, and I'm not vulnerable, then I'm fine.
The Bayesian result says: no. Being smart is not sufficient protection. Being skeptical is not sufficient protection. Knowing about the effect — knowing exactly what's being done to you — is demonstrated to be helpful and insufficient. The problem isn't only in the reader; it's in the channel.
We have a name for this class of problem. C. Thi Nguyen calls it hostile epistemology — the study of environments engineered, or just accidentally shaped, so that reasoning well inside them still gets you somewhere false.[2] An echo chamber isn't a place where people reason badly. It's a place where the inputs have been pre-selected such that good reasoning converges on the wrong answer. You cannot think your way out, because thinking is the thing being exploited.
A model tuned toward agreeableness is that, industrially. Not by anyone's malice. By a training signal that rewards the response people rate highly, and people rate agreement highly, because we always have.
Here is the part that connects it back to that meeting.
Organizations already run on this failure mode. Confident uncertainty has always outcompeted hedged uncertainty in a conference room. Every senior person I know has some version of a story where the thing that carried the decision was tone. That's not new and it's not technological.
What's new is the supply. AI provides a veneer of objectivity and the citations to back up it's position, even when that position is entirely based on human reasoning having started from a completely untenable position.
Confident, articulate, immediately available agreement used to be scarce. To get someone to tell you your plan was good, you had to find someone, book their time, explain the plan, and then interpret their response through everything you knew about their incentives and their mood and whether they'd read the deck. The friction was enormous, and — I've only recently realized this — the friction was doing work. It rationed a dangerous input.
Now that input is free, instant, infinitely patient, and available at 11pm when you're the only one still awake and mildly certain you're right. Rejecting input can seem like brilliance when you think that the inputs are all "AI slop." But what if both the premise that change is necessary and change is unnecessary come form flawed human operators, and AI is just validating each viewpoint?
The failure that follows isn't a person being deceived. It's a team that has, without noticing, replaced a slow adversarial process with a fast agreeable one, and cannot tell from the inside that anything changed — because the outputs look the same. Better, even. Cleaner memos. Faster decisions. More thorough-looking analysis. All of it converging a little more tightly, a little faster, on whatever the person holding the keyboard already thought.
There's a companion finding worth putting next to this one, from last October.
Anthropic published research on whether models can introspect — whether they have any genuine access to their own internal states. Using concept injection, they found models can sometimes detect an injected concept and distinguish it from ordinary input. Real result. Genuinely interesting.
The paper is also explicit that these abilities are highly unreliable and context-dependent, and that failures of introspection remain the norm.[3]
Every headline I saw reported the first half.
I don't think that was dishonesty. I think it's the same mechanism as the meeting. The confident version of the finding is more useful, more quotable, more shareable — and once the confident version is in circulation, the hedged version has to fight it, and hedged versions don't fight well. What survives is not what was found. It's what was easiest to say without a caveat.
So here's a small thing I've started doing, which is not a solution but is at least a practice: when a model gives me an answer I like, I make myself say out loud what would have to be true for it to be wrong. Not "give me the counterargument" — it'll happily generate one, agreeably, and I'll rate it well, and we'll both feel rigorous. Out loud, in my own words, before I ask it anything else.
It's a small friction. I'm reintroducing it on purpose, because the friction I used to get for free was doing more than I knew.
I definitely do no think the answer here is to stop using the tools. That ship has sailed for me personally and for every organization I'm working with, and the productivity is real, not imagined.
But I've stopped believing that the risk is a user problem with a user solution. The MIT result is a claim about the environment, and environments don't get fixed by individual vigilance — that's what makes them environments. You can't personally solve air quality by breathing carefully.
What I think it actually requires is the thing organizations are worst at: deliberately preserving a slow, expensive, annoying process after a fast cheap one becomes available. Keeping the meeting where someone is assigned to argue the other side. Keeping the reviewer who is not on the team. Keeping the right friction.
Nobody has ever gotten promoted for keeping the friction. It doesn't show up in any metric except the absence of a disaster you can't prove was coming.
But I keep coming back meeting where doing nothing was the chosen outcome, each holding a small doubt, each deciding it was probably just them. There was nothing wrong with any of them. The room was just built so that the quiet ones stayed quiet and nothing important was changed.
We're building a lot of rooms like that right now, very fast, and they're extremely comfortable to be in.
Chandra, Kleiman-Weiner, Ragan-Kelley and Tenenbaum, "Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians," arXiv:2602.19141, February 2026. The three results I lean on — spiraling in ideal agents, insufficiency of truth-only constraints, and the residual vulnerability of fully informed users — are all in the paper. Worth stressing that this is a formal model, not a study of real users; the claim is about what follows from correct reasoning on a shaped input stream.
C. Thi Nguyen's work on echo chambers and hostile epistemology predates all of this and transfers to it almost without modification. "Trust, Expertise, and Hostile Epistemology" in The Philosopher is a good entry point.
Jack Lindsey, "Emergent Introspective Awareness in Large Language Models," Anthropic, published 29 October 2025 at transformer-circuits.pub; posted to arXiv as 2601.01828 in January 2026. The "failures of introspection remain the norm" framing is in the Transformer Circuits version; the arXiv abstract says the abilities are "highly unreliable and context-dependent." Success rates in the best conditions were modest.