Language models
A chat tool's reply is one prediction, made again and again. See how a language model predicts the next word, how it learns to, how a whole answer grows, and how training turns it into an assistant.
A language model writes its answer one small piece at a time. Each piece is a prediction of what comes next, based on everything written so far. This module takes that one idea and follows it all the way to the assistant you chat with.
One prediction¶
Pick a prompt and step through it. The text is split into tokens, small pieces of words. The model weighs every earlier token to decide what matters for the next one, blends them into a single point, and scores every word it knows by how close that word sits to the point.
Run the chocolate prompt, then the sock prompt. Both sentences end with the vet about to act, yet one leads to vomiting and the other to radiographs. The model reads the whole sentence, not just the last word.
How it learns¶
A model learns to predict by practising on text. Every position in a sentence is a free example: given the words so far, guess the next one, then check the answer. Train one pass on "Chocolate is toxic to dogs." and watch each guess get scored. Then fast-forward and watch the right words climb from random guesses to near certainty as the loss falls.
This toy memorises one sentence, because each position has its own scores. A real model shares its weights across everything it reads, so it has to find patterns that hold in general.
A whole reply¶
A reply is the same single prediction run in a loop. Each new token is added to the input, and the model predicts again. Compare one next word with a full response, and watch the counter: one forward pass for every token.
Then raise the temperature and generate a few times. A higher temperature gives less likely words more of a chance. Now and then the first word is "Yes", and the rest of the answer follows it, fluent and wrong. One bad early word steers everything after it.
From predictor to assistant¶
A model trained only to predict text does not answer questions. It continues the page. Step through the three stages with the same question, "Can my dog eat grapes?". The pretrained model writes what tends to follow such a question on the web, such as more headlines. After instruction tuning, it answers.
At stage 3 you are the rater. Pick the warm, friendly answer and see what the model learns: it says grapes are usually fine. A model trained to please can learn to be wrong, which is one reason chat tools tend to agree with whoever is asking.
What to take away¶
Everything a chat tool writes comes from next-word predictions, repeated and then shaped by training to sound helpful. Sounding right and being right are different things, and that is why you check what it tells you.
Where next
How AI works: Module 3: Retrieval
Back to your lesson: Using AI in practice, Module 1: How LLMs work
A shared question channel is on the way. When it opens, each answer will be written once and shared with everyone taking the course.