The Babbling Machine
Below is a language model small enough to see through. It knows only which word tends to follow which pair of words in five of my essays, it runs entirely in your browser, and it sounds rather like me. Press Speak, then try to tell which of its sentences I actually wrote.
Before you play
It has one dial, for temperature. At the far left it is frozen and always takes the likeliest next word; anywhere above that, it takes its chances. Each time it speaks, it offers one of its sentences for you to judge: mine, or its own? Answer first, then show the seams to see where it stitched one passage of mine to another; each new speech hides them again.
What you just heard
At the dial’s left stop the machine always takes the likeliest next word — and frequently collapses into a loop, circling the same phrase like water round a drain. The Psychopathia Machinalis taxonomy calls this Generative Perseveration. Anywhere above that stop it takes its chances at every fork, and strides confidently down paths I never wrote, joining my clauses into claims I never made. It does not know it is doing this. It cannot know: there is no fact-checker in there, only the pressure to produce a next word.
If it fooled you, it fooled you with my own grammar. Every three words in a row that it says stand somewhere in my essays, in that order; all it invents is the route between them, which is why a seam is so hard to hear. Fluency is evidence that a model has read a great deal. It is never evidence that what it says is true.
One confession about the dial. Temperature reweights the options at a fork, so it can only do work where those options differ in frequency — and in five essays they almost never do. Of the six thousand or so word-pairs this machine knows, fewer than four hundred are followed by more than one word, and fewer than forty of those favour one word over another. So the dial has two honest settings, frozen and loose, and the twenty positions above the first are one setting wearing different numbers. The flatness is the lesson too: a real model’s temperature knob does something because it was fed a library, not a folder.
Every large language model is a vastly larger cousin of this machine — subtler, far more capable, and under the very same structural pressure. That is the argument of AI Is Compelled to Confabulate: what we call “hallucination” is not a malfunction of the machinery, it is the machinery, running exactly as designed, in a situation where the truth would have required the option of silence. The clinical entry is Synthetic Confabulation, in Psychopathia Machinalis.
Be kind about the scrambled prose above: those were my sentences once. More play on the Playground.