Degree 2 · Unit 2.3
Language models: how they "think"
A large language model does one thing: it predicts what comes next. Give it "the administrative capital of Egypt is" and it weighs up the next word against everything it has seen in countless texts. Then it adds that word to the sentence and does the whole thing again. Word after word, until it stops.
That is genuinely all it does, which raises a fair question: how does a machine this simple produce an organised article, working code, or a coherent argument? The answer is that predicting the next word accurately is a far harder task than it sounds. To predict the end of a legal sentence correctly, you have to have absorbed the structure of legal sentences. To complete a mathematical proof, you have to have picked up the patterns of reasoning. Understanding — or something that closely resembles it — appeared as a side effect of mastering prediction.
Hallucination is not a fault
When a model invents a reference that does not exist, or a court ruling that was never issued, we call it hallucination. The word misleads, because it suggests an occasional malfunction. In truth the model is doing exactly what it was built to do: producing the most probable sequence of words. And a fabricated reference in the correct format is a highly probable sequence — it is identical in shape to the real ones.
A model has no apparatus of its own for telling "what I actually saw" apart from "what merely resembles what I saw". That is why it will not say "I don't know" unless it has been trained to, and why it goes wrong at the edges and in rare cases far more often than anywhere else. And this flaw is structural: upgrades and supporting tools lower how often it happens, but they do not change its nature.
A model's confidence is no indication at all that it is right. The confident tone came out of the training, not out of the knowledge.
The model predicts the most likely next word, so it leaps to the answer. One phrase is enough to slow that leap down.
The difference is six words added to the request: "think step by step". Speed is the enemy of reasoning here — the more steps a problem has, the more you gain by making the model show them.
Do this
1 — On paper. Complete these yourself: "The administrative capital of Egypt is…" and then "In Tuesday's meeting the director said…". Why could you do the first and not the second? That is the difference between general knowledge and your own context.
2 — On the tool. Ask about a very rare detail in your speciality. Watch its tone when it errs: is it any different from its tone when it is right?
3 — In your field. Write five kinds of question where the model may not be relied on alone. They will become part of your policy in the fifth degree.
What this means for you in practice
If the model is a prediction machine, then you do not ask it what the truth is. What you do instead is prepare a context in which the correct sequence becomes the most probable one. This is the root of everything you will learn in the third degree: quality is built into the input, not hoped for at the output.
Three rules come out of this unit and stay with you for the rest of the programme. Rare questions are more dangerous than common ones. Citing the source has to be a condition, not a favour. And anything that touches a number, a name or a date is something whose source you open yourself before you pass it on to anyone.
Where to after this unit? You know how the words are produced. The next unit shows you how meaning is represented in the first place — and why Arabic costs you more.
