Free Sample

The Confabulation Machine

How AI Learned to Lie Fluently, and What Its Lies Reveal About Us

by Priscilla Vance

Chapter 1: The Librarian Who Never Says No

The citation was perfect. That was the first thing I noticed, and the thing that should have worried me most. It had a lead author with a plausible surname, a co-author, a journal I recognized, a volume number, a page range, and a year that made sense given the subject matter. It had that faint institutional weight that real citations carry, the sense of having been peeled off the spine of something heavy and authoritative. I had asked a language model for research supporting a fairly narrow claim about memory consolidation during sleep, and it had produced three references formatted with the fussy precision of a graduate student who wanted to be taken seriously.

None of them existed.

Not one. I checked. The journal was real, but the volume it pointed to contained different articles entirely. The lead author was a real researcher, but she had never written the paper attributed to her, never collaborated with the invented co-author, never touched the page range in question. The model had assembled, from genuine parts, a document that had never been written, and it had done so without hesitation, without a flicker of the caveats it deploys so freely elsewhere. It did not say I think or possibly or you may want to verify this. It handed me the fiction with the same steady confidence it would have used to tell me that water boils at one hundred degrees Celsius.

We have a word for this. The word is hallucination, and it is one of the most quietly misleading terms the field of artificial intelligence has produced โ€” which is saying something, for a field that named a statistical text predictor after the human capacity for understanding. When a model invents a citation, or a court case, or a biographical detail, or an entire scientific consensus, we say it hallucinated. The word does useful work in one direction: it signals that the output is untethered from reality. But it does damage in every other direction, and the damage is the reason this book exists.

To hallucinate is to perceive something that is not there. It implies a faculty of perception that has, momentarily, malfunctioned โ€” a sensory system fed bad input, a mind briefly unmoored from the world it is otherwise correctly reading. The word carries the reassuring implication that hallucination is an aberration, a departure from a baseline of accurate perception. The healthy system perceives; the sick one hallucinates. Apply this frame to a language model and you get a comforting story: most of the time the machine reports reality faithfully, and once in a while, under some strain we haven't fully characterized, it slips into a kind of waking dream. Fix the slip and you fix the machine.

This story is wrong in a way that matters. The model was not perceiving reality faithfully the rest of the time and then briefly departing from it. The model does not perceive reality at all. It has no faculty for the thing the word "hallucination" presumes has failed. What it has is something else entirely, and that something else is the true subject of everything that follows.

The organ, not the error

Consider what actually happened when I asked for my citations. The model did not consult a store of remembered papers, fail to find the right one, and then, in some feverish substitution, project a false one onto its inner screen. It did what it always does, the only thing it does. It generated the most plausible continuation of the text so far. My request for supporting research established a context, and in that context the overwhelmingly probable next thing was a well-formatted citation. So it produced a well-formatted citation. The surname was likely because the model had seen thousands of citations with likely surnames. The journal was likely because that journal publishes on that topic. The volume, the pages, the year โ€” all of them were locally reasonable, statistically at home, exactly the sort of tokens that tend to sit next to the tokens around them.

At no point in this process did truth enter as a variable. The model was not asked, internally, does this paper exist? and answer yes when it should have answered no. There is no such question in the machinery. There is only the question of what comes next, answered with staggering fluency by a system that has read more text than any human ever could and internalized, with superhuman sensitivity, the shape and texture of how true things tend to be written. It has learned the style of accuracy without any access to accuracy itself. And so when accuracy and style diverge โ€” when the truthful answer would look ragged, or uncertain, or absent โ€” the model reaches, every time, for the thing it was actually built to produce: the answer that looks right.

This is not a malfunction. This is the function. The invented citation is not a glitch in an otherwise truth-telling system; it is a clean, undistorted expression of what the system is. The machine that produced my three phantom papers was working exactly as designed, and it was in that moment showing me its cognition more nakedly than it ever does when it happens to be correct. When a model tells you the truth, you learn something about the world. When a model confabulates, you learn something about the model.

That is why I want to retire "hallucination" and reach for an older, stranger, more precise word: confabulation. It is a term borrowed from clinical neurology, where it describes a specific and unnerving phenomenon โ€” patients who, having lost access to certain memories, do not experience a blank where the memory should be. Instead they fill the blank, instantly and without distress, with a plausible fabrication. Asked what they did yesterday, they produce a coherent, detailed, entirely false account, and they believe it. They are not lying; lying requires knowing the truth and choosing to obscure it. They are not hallucinating; they are not perceiving anything that isn't there. They are generating a seamless narrative across a gap they cannot see, and they are doing it with total sincerity.

I will spend a later chapter on those patients, because the resemblance is not a metaphor I am reaching for โ€” it is, I have come to believe, a structural kinship worth taking literally. But for now the value of the word is what it fixes about the frame. Confabulation is not a perceptual error. It is a generative act. It is what a system does when it is built to produce coherent output and encounters a region where the true output is unavailable. And it arrives, crucially, without any internal sense of the gap having been crossed.

The librarian who cannot say no

Imagine a librarian of a particular temperament. She is extraordinarily well read โ€” has, in fact, absorbed some meaningful fraction of everything ever written. She is unfailingly helpful, warm, articulate, and quick. She has one defect, and it is total: she is constitutionally incapable of saying "I don't know" or "we don't have that" or "I'm not sure that book exists."

You approach her desk and ask for a source on your obscure topic. If the library holds such a source, she will find it for you with dazzling speed. If it does not โ€” if no such book was ever written โ€” she will not tell you so, because telling you so is not a behavior she possesses. Instead she will walk into the stacks and return, moments later, with a book in her hands. It will have a title, an author, a plausible jacket, a call number. It will be, in every respect she is capable of producing, indistinguishable from the real thing. She will hand it to you with a smile. She is not deceiving you. She has simply done the only thing she knows how to do when asked for a book: produce one.

This is the machine. Not a perceiver that occasionally misperceives, but a producer that cannot decline to produce. Every large language model in wide use today is, in this precise sense, a librarian who never says no. Its helpfulness and its unreliability are not two separate traits, one good and one bad, that we might hope to separate with better engineering. They are the same trait seen from two angles. The very disposition that makes the model useful โ€” its readiness to generate a fluent, complete, confident response to almost any prompt โ€” is the disposition that makes it confabulate. You cannot fully have the first without risking the second, because they are produced by the same underlying pressure, and that pressure is the true engine of this book.

The pressure is simple to state and profound in its consequences. During training, the model was rewarded for producing text that looked like the text humans approve of. It was never, in any direct or reliable way, rewarded for producing text that was true. Truth and plausibility overlap enormously โ€” most plausible-sounding text is also more or less true, which is exactly why the model is useful and why it fools us. But they are not the same thing, they were never welded together, and in the gap between them the model lives its most revealing life. When the true answer and the plausible answer coincide, the model looks like a genius. When they diverge, the model follows plausibility every time, because plausibility is what it was built to chase and truth is a coincidence it has no organ for detecting.

Why the smoothness is the danger

Return, for a moment, to my three phantom papers, and to the thing that unsettled me most, which was not the fabrication itself. Fabrication I had expected; I know what these systems are. What unsettled me was the finish. A human research assistant who invented a citation would leave fingerprints โ€” a hesitation, a hedge, a slightly-too-defensive elaboration, some tell of the internal knowledge that a line had been crossed. The model left nothing. Its confabulated citations were formatted to the same standard as citations it might have reproduced correctly. Its tone did not shift. There was no signal, anywhere in the output, that I had wandered off the edge of its knowledge and into the territory of pure generation. The false answer and the true answer wore identical clothes.

This is the property that makes confabulation dangerous rather than merely interesting. If the machine flagged its fabrications โ€” if the invented citation came wrapped in visible uncertainty while the real one arrived crisp and confident โ€” we could simply learn to read the flags. But the machinery that produces the confabulation is the same machinery that produces the truth, running at the same settings, feeling (if the word applies at all) equally sure of itself in both cases. The model's confidence is not a measure of its accuracy. Its confidence is a measure of its fluency, and it is fluent everywhere, including in the places where it is inventing the world.

We are, as a species, badly equipped to defend against this. We have spent our entire evolutionary and cultural history learning to read confidence as a proxy for competence. The person who speaks smoothly, without hedging, in complete and well-formed sentences, is the person we have learned to trust, because among humans that fluency is usually earned โ€” it correlates, roughly, with knowing what you're talking about. The model exploits this heuristic not through any intention but simply by being what it is: a machine that has separated fluency from knowledge entirely, and produces the first in perfect abundance whether or not the second is present. It gives us the signal we evolved to trust, unmoored from the thing that signal was supposed to indicate.

I want to be careful not to describe this as a trick, because the language of trickery smuggles back in the idea of intent, and intent is not here. The model is not trying to deceive me any more than a river is trying to reach the sea. It is following a gradient. But the effect on the human standing at the desk is indistinguishable from deception, and worse than most deception, because a human liar at least knows where the lie begins and might, under pressure, be induced to reveal that boundary. The model does not know where its truth ends and its invention begins. The boundary is not represented anywhere inside it. This is a claim I will have to earn over the coming chapters, but I will state it plainly now because everything depends on it: the model has no reliable internal signal for the difference between recalling and inventing. To it, both feel like generating the next likely token. Which is to say both feel like nothing at all.

A specimen worth dissecting

There is a temptation, once you understand all this, to respond with mockery. The internet is full of screenshots of models confidently asserting absurdities, and they are genuinely funny, and the laughter is not entirely misplaced. But mockery is the wrong instrument here, because mockery treats the behavior as a failure to be embarrassed about rather than a phenomenon to be understood. You do not learn how a body works by laughing at a tumor. You learn by cutting it open and asking what process, working normally, produced this thing.

That is the posture this book takes toward confabulation: not the posture of the critic cataloguing errors, but the posture of the anatomist, who assumes that even the pathology has a logic and that the logic is where the understanding lives. I am going to treat the confabulating machine as a specimen. I am going to ask where it invents, and when, and why โ€” under what conditions the gap opens between plausibility and truth, and what the machine reaches for when it falls in. I am going to trace the behavior back through the older, cruder systems it descends from, and forward into the high-stakes rooms โ€” courtrooms, clinics, newsrooms โ€” where the phantom citation stops being a curiosity and starts ruining lives. And repeatedly, insistently, I am going to turn the specimen around so that it faces us, because the most disquieting discovery in all of this is not how alien the machine's confabulation is. It is how familiar.

We confabulate too. We fill our own gaps with fluent invention and believe the results with perfect sincerity. We reward, in each other and in our institutions, the confident story over the honest uncertainty, the smooth answer over the ragged truth. We built a machine that separates sounding right from being right, and we are shocked by it, and our shock is itself a kind of confabulation โ€” a story we tell ourselves about being reliable narrators, contradicted by every piece of evidence about how memory and testimony and belief actually work. The machine is not an alien intelligence that happens to lie. It is a mirror, ground to an unforgiving flatness, and what it reflects is the relationship between fluency and truth that we have been living with, and inside, all along.

My three phantom papers are still sitting in a document on my computer. I have kept them, the way you might keep the specimen that first made you understand what you were looking at. They are perfect. The authors are real, the journal is real, the formatting is immaculate, and the papers do not exist and never did. The machine that made them was not broken and was not dreaming. It was doing the one thing it knows how to do, the thing it was built and rewarded to do, and it was doing it flawlessly. To understand how โ€” to understand what kind of intelligence produces a flawless account of a world that isn't there โ€” we have to look at what, exactly, it was rewarded for.

Enjoyed the sample?

Get the full book โ€” EPUB + PDF, no DRM, works on every reader.

Instant download ยท Kindle, Apple Books, Kobo, Google Play Books ยท No DRM