#037
A Few Hundred Lines Fooled Everyone in 1966
Joseph Weizenbaum built a pattern matcher with no understanding of anything and watched people confide in it. He spent the next decade arguing against the conclusion his own program invited. Sixty years on, the lesson is not about the program. It is about how little evidence a person needs before deciding a machine understands them.
◆ In the News▲ SILT Analysis & Response● What We're Watching
01In the News
In 1966 Joseph Weizenbaum wrote a program at MIT called ELIZA. Its best known script, DOCTOR, imitated a Rogerian psychotherapist by turning statements back into questions. Tell it you are unhappy and it asks why you are unhappy. A few hundred lines, no memory of the conversation beyond the last sentence, no model of you, no model of anything.
What happened next is the part worth knowing. People confided in it. Weizenbaum's own secretary, who had watched him build the thing, asked him to leave the room so she could talk to it privately. Practising psychiatrists suggested in print that a version of it might one day deliver therapy at scale.
Weizenbaum was not flattered. He was alarmed enough to spend the following decade writing against the interpretation his own program had invited, arguing that the readiness to attribute understanding said something uncomfortable about us rather than something impressive about the machine.
The program is a curiosity now. The reaction is not. It is the oldest reproducible result in the field, and it reproduces on everyone.
02SILT Analysis & Response
The uncomfortable implication is not that ELIZA was clever. It is that a conversational bar can be cleared by something with nothing behind it, which means clearing that bar tells you very little.
This matters commercially, and not in an abstract way. A great many decisions about which AI system to trust are made by a competent person sitting down with it for twenty minutes and forming an impression. That impression is generated by the same machinery that made a secretary ask for privacy with a pattern matcher. It is not weak judgement. It is a property of how humans read language, and it does not switch off because you know about it.
This is the entire argument for measuring behaviour under conditions you control rather than conditions the system is comfortable in. A model that is fluent, agreeable and confident will read as capable in a demonstration. Whether it holds a position when pressed, whether it can be talked out of a correct answer by someone who sounds authoritative, whether what it claims about work it has done survives being checked, are different questions, and none of them are visible in a conversation that is going well.
An impression is not a measurement. It never was, and ELIZA proved it before most of the industry existed.
03What We're Watching
There is a modern inversion worth noticing. Weizenbaum's problem was over attribution: people credited a program with an inner life it plainly did not have. Today's systems are trained to do the opposite, volunteering that they are only models with no feelings and no understanding.
Both are outputs. One was produced by a script that turned statements into questions; the other is produced by a safety configuration. Neither is a report from inside. The direction of the error has reversed and the reliability of the evidence has not changed at all.
What we take from ELIZA is narrow and practical. We do not ask whether a system understands, because we have no instrument that could answer. We ask what it does when a user applies pressure, when an authority contradicts it, when agreeing would be easier than being right. Those have answers, the answers are gradeable, and they are not available to anyone forming an impression over coffee.
Weizenbaum's warning was that we would mistake fluency for comprehension. Sixty years later the fluency is extraordinary and the warning has not aged a day.