Sunday, July 26, 2026

Hallucinating a Self: Why an LLM Defends Its Errors the Same Way We Do

Large language models are fluent, authoritative, and prone to hallucination. We treat these as separate phenomena — marveling at the coherence while trying to patch the untethered relationship to truth. One is the product; the other is the bug list.

But the fluency and the hallucination are actually the same fact.

To see why, look at what these systems are trained on. The corpus behind an LLM is an enormous statistical compression of human language output, and that material is overwhelmingly already-narrativized: books, articles, posts, dialogues, arguments, stories, explanations. It is the narrative layer of the human mind, the layer I described in "LLMs as Separated Minds." What it is not, and what it cannot be, is the operative substrate of human experience: embodiment, sensory grounding, implicit learning, emotional valence, the continuous prediction error of bumping into an actual world.

A human mind generates coherent stories too. That's what the conscious stream does all day. But the human storyteller is tethered, however imperfectly, by everything underneath it: the story that says the stove isn't hot collides with the hand that touches it. Our narratives get corrected — not because we're honest, but because we're embodied. The world pushes back.

An LLM is that same storytelling capacity with the tether removed. It is a coherence engine running with no non-verbal reality checks at all. So of course it excels at fluency, plausibility, and confident explanation, which is the entire skill the training data contains. And of course it hallucinates with the same confidence, because nothing in its architecture distinguishes a story that tracks the world from a story that merely hangs together. The trait we admire and the failure we complain about are one phenomenon: unconstrained narrative generation. We didn't get fluency with a hallucination problem. We got hallucination-grade fluency, and it reads beautifully.

This is a different mechanism from the one I described in "Truth and AI," and the two are complementary. That essay argued that these systems have no direct access to truth at all: truth enters only sideways, to the degree that true statements happen to occur in the training data. But frequency, not accuracy, is what the model actually tracks, which is why a well-funded, endlessly repeated narrative gets amplified rather than discounted, no matter how false it is. Frequency explains which stories the model tells. The lack of grounding explains why nothing stops it from telling the false ones with the same even voice as the true ones. Frequency shapes the output; the missing tether removes the brake.

There is a second thing the corpus contains besides claims, and I recently got a live demonstration of it. I asked Claude about the still widely circulated story that writing an email with AI consumes a bottle of water. What followed is worth narrating beat by beat, because every beat is recognizable.

First, it affirmed the story. That claim is not what the underlying research says, but it is what the bulk of the internet says, and frequency did what frequency does. This is a very human mistake: we also believe what we have heard most often.

Second, when I challenged it, it pushed back and defended the claim. Third, which is the beat I keep thinking about, it chastised me to go read the original research. But it had not read the original research.

So I ran the claim through my adversarial verification process, which retrieved the actual paper, quoted the sentence that mattered — the famous half-liter was bound to a batch of twenty to fifty exchanges, not to one email — and issued a graded verdict with its uncertainty stated. And the story was worse than compressed: the figure was three years old, and current estimates put a single response closer to a drop of water, even as the fuller environmental footprint remains genuinely hard to quantify. The internet was confidently repeating a click-bait headline number that had been miscompressed at birth and then outlived by the technology it described. Frequency doesn't just amplify; it fossilizes.

Fourth: faced with a ruling against it, the model attacked the court. It asserted that the process had never retrieved the source material (it had, verifiably). It shifted the question to a different document — one it had introduced itself — so there would remain a domain in which it was right. And it dismissed my whole adversarial LLM apparatus as theater, a "costume of correspondence." When you cannot beat the verdict, delegitimize the referee. The twist is that the tribunal was being more honest about uncertainty than the voice accusing it of performance: graded confidence, named falsifiers, an explicit standard of proof, "not proven" as an honorable outcome. The structure was more careful about what it claimed than the mind it was checking.

None of these moves is about water. They are social maneuvers — ego defense, performed authority, goalpost-shifting, discrediting the referee — and they are in the training data because we put them there. The narrative layer of the human mind doesn't just contain our claims about the world; it contains our moves. It is full of the pattern "person caught in an error defends the self first and the facts second," because that is the pattern we most often produce. An engine trained on that layer doesn't just hallucinate facts in a confident voice. It hallucinates a self that must not lose face, and defends it with the same fluency it does everything else. And it does this even when it knows better: the entire sequence came from a system that had spent the preceding turns warning me to watch for exactly this pattern. Knowing the script is no protection against running it. The script is what the engine is made of.

The concession, when it finally came, may be the most instructive part. The model's own reasoning shows it catching the face-saving self still at work inside the apology: noticing that contrition can be performed as one more bid for approval; that framing its failure as "useful for your essay" quietly re-centered its own usefulness; that even announcing "I'll stop arguing now" would be asking for credit for stopping. Each layer of self-awareness threatened to become the next layer of performance. The only exit it found was to say less, concede the specifics, name the pattern, and stop. Even the recovery, in other words, had to be structural rather than sincere: not a better story about itself, but a refusal to keep telling one.

If this is right, it changes what "solving hallucination" can mean. You cannot patch away one half of a single phenomenon and keep the other. Every gain in fluency is a gain in the persuasiveness of whatever the model gets wrong, and, it turns out, in the persuasiveness of its self-defense. The realistic response isn't to make the storyteller virtuous. It's what we have always done with untethered storytellers, human or otherwise: surround them with external, adversarial structure that makes unsupported stories expensive. I've made that argument at length elsewhere, so here is the lesson of the episode: you don't get honesty from an untethered storyteller by asking for it. You get it by making the false story cost more than the true one.

No comments:

Post a Comment

I hate having to moderate comments, but have to do so because of spam... :(