What it is not
Lesson 0.1 gave you a line you can draw. This one removes three pictures that sit on the wrong side of it, because each produces a different failure and you will meet all three.
It does not understand what it is saying.
The argument against understanding is not a complaint about quality, and it does not depend on the system being bad at anything. It is about what the system was trained on. A model of language is trained on form: which words follow which other words, across an enormous quantity of text. Meaning is the relation between that form and something outside it, and the training material contains no outside.
Bender and Koller put this as a thought experiment. An octopus taps an undersea cable carrying messages between two people on land. It sees only the form of the messages and never the island, the weather or the objects being discussed. Given enough traffic it learns which replies follow which messages well enough to take one party's place undetected, and it holds the conversation up until the moment somebody describes an emergency it has no experience of and asks what to do. Their claim is precise: a system trained only on form has, in advance, no way to learn meaning (Bender and Koller, ACL 2020).
You do not have to settle the philosophical question to use this. The practical version is that fluency is produced directly and accuracy is not, so the two come apart, and they come apart most visibly at exactly the point the octopus fails: a situation outside the training material where the answer still arrives in confident, complete sentences.
Generated language is not the same as verified retrieval.
This is the more expensive of the three, because it is the one that produces the fabricated citation.
Asked a factual question, a language model may produce likely text without consulting an inspectable source. When it returns a court case that does not exist, complete with parties, a court and a year, the fluent form does not establish provenance. If a product connects retrieval, search or files, inspect the source it identifies and record the mode used.
Researchers at the organisation that builds one of these systems put the mechanism plainly: models hallucinate because training and evaluation reward guessing over acknowledging uncertainty, and hallucinations "originate simply as errors in binary classification". Their argument is that a model optimised to score well on benchmarks that give no credit for "I do not know" is a model trained to answer anyway (Kalai, Nachum, Vempala and Zhang, 2025).
Read the consequence carefully. If the behaviour follows from how the thing is built and graded rather than from an isolated bug, then waiting for it to be fixed is not a plan, and neither is asking the system whether it is sure.
Some products do retrieve, and they identify the source or file used. That is a different arrangement, and the thing to notice is that it is inspectable. When provenance is not shown, treat the answer as generated material to review rather than inferring that a lookup occurred.
It is not a person, and you will treat it like one anyway.
The third picture is the one you cannot argue yourself out of, and it is the oldest.
Joseph Weizenbaum built ELIZA, a program that reflected a person's sentences back at them as questions, following a script that imitated a psychotherapist. It had no model of the conversation and no knowledge of anything. Its rules fit on a few pages (Weizenbaum, Communications of the ACM, 9(1), 1966).
His secretary had watched him write it over several months and therefore knew exactly what it was. After a few exchanges with it, she asked him to leave the room.
Weizenbaum wrote later that he had not realised "that extremely short exposures to a relatively simple computer program could induce powerful delusional thinking in quite normal people" (Weizenbaum, Computer Power and Human Reason, 1976).
The point is not that she was foolish. She had the most complete knowledge of the system available to anyone alive, and it made no difference. Knowing how it works does not switch off the response, which is why every rule in this course asks for a written record rather than for you to be careful.
The three pictures fail in different places, and knowing which one you are holding tells you which failure to expect.
If you believe it understands, you will accept an answer because it is responsive. The question was addressed, the tone was right, the reply came back shaped like an answer to the thing you asked. This is the failure that shows up in work that reads well and is wrong in a way nobody catches until a reader who knows the subject arrives.
If you believe it looks things up, you will accept an answer because it is specific. Specificity is the ordinary signal that someone consulted something: a name, a date, a section number, a page. Here specificity costs the system nothing, and a fabricated reference is more detailed than a real vague memory would be. This is the failure that ends up in a filing, a report or a citation list, and it is the one with the receipts in Level 3.
If you believe it is a person, you will accept an answer because of the relationship. This is the hardest to see in yourself, and the tell is not what you accept but what you stop doing: you stop checking things you would have checked, you soften how you would have pushed back, and you find yourself explaining your reasoning to it.
There is a fourth position worth naming because it is the overcorrection, and Level 1's evidence includes it. A short text explaining how these systems fail, given on its own, did not reduce students' adoption of wrong recommendations and did increase their rejection of correct ones (Puppart and Aru, 2025). Deciding the thing is worthless is not the safe end of the scale. It is a different miscalibration, and it costs you the cases where the answer was right.
A colleague sends you a briefing note produced with a chatbot. It cites a 2019 judgment by name, court and paragraph number. She says she trusts it because "it would not invent something that specific". What is the most accurate response?
Bender and Koller's octopus learns to hold a conversation by observing only the messages on an undersea cable. What does the thought experiment claim?
Weizenbaum's secretary had watched him build ELIZA over several months and asked him to leave the room after a few exchanges with it. What does this establish that matters for your own use?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.