Phones have done this for years: you type "I'll be there in five" and the keyboard offers "minutes." A large language model is, at heart, the same idea scaled up enormously — a system that looks at the text so far and predicts what is likely to come next. That sounds almost too simple to power something that writes essays and code, but the surprise of the last few years is that when you make that predictor big enough and train it on enough text, it gets startlingly good at it.
It predicts the next token
Models don't actually work in whole words; they work in tokens — small chunks of text, often a word or a piece of one. A common word like "cat" is usually a single token, often with the space in front of it included, while a rarer word gets split into pieces: "unbelievable" might become "un", "believ" and "able". As a rough rule of thumb, a token is about three-quarters of an English word. "Predicting" means: given all the tokens it has seen so far, the model estimates which token is most likely to come next, picks one, and adds it to the text. There is no grand plan or sentence laid out in advance. The reason a model can answer a question or write a function is that, across a vast amount of training text, the most likely continuation of a well-posed question is its answer. Good predictions, stacked one after another, look a lot like thinking.
How it works
Everything you send — your question, any instructions, any files — becomes a sequence of tokens that the model reads all at once. It then gives every token in its vocabulary a probability of coming next, draws one according to those odds, appends it, and repeats: read everything, score, draw, append, read again. That loop is why output streams out word by word and always reads left to right — each new token is generated with full sight of everything before it but nothing after.
The draw is where temperature comes in. It's a setting that reshapes the odds before the pick. A low temperature sharpens them, so the favourite almost always wins and the same prompt gives nearly the same answer every time. A high temperature flattens them, so less likely tokens get picked more often and answers vary more. Neither setting adds knowledge; it only changes how adventurous the draw is.
Step through the loop below, and flip the temperature switch to watch the odds sharpen and flatten. In the last part the model is asked about a name it has never seen: predict what it writes before you look.
In our stack — the models doing this prediction are Anthropic's Claude models. When Claude Code works on your project, it bundles your request and the relevant code into tokens, sends them to a Claude model, and streams back the predicted tokens — which might be an explanation, a patch, or a decision to call a tool. The model itself stays the same between requests; all the project-specific knowledge rides along in what gets sent.
You send the exact same prompt twice and get two differently worded answers. What's the most likely reason?
Fluent isn't the same as true
The end of that scene is the most important thing to understand about these models. The loop that wrote "mat" is the same loop that wrote a function name that doesn't exist. At every step the model must produce some token, and it picks from what is likely, not from what has been checked. When the true answer was common in its training text, likely and true line up. When it wasn't — your private codebase, the newest version of a niche library, a paper nobody wrote — the model still produces the most plausible-sounding continuation, in exactly the same confident tone. This is what people call a hallucination.
For developers it shows up in predictable places: invented API methods and config flags, package names that were never published, version numbers and release dates, citations and links. Newer models are trained to say when they're unsure, which helps, but it reduces the problem rather than removing it. The dependable fix has two halves: put the real facts in front of the model (the actual source file, the current docs), and check what comes back, ideally by running it.
Confidence is not evidence. A model's tone tells you nothing about whether it's right: a made-up method name reads exactly as smoothly as a real one. Turning the temperature down makes answers repeatable, not correct. Treat anything that names a specific API, version, file or fact as a claim to verify, and hand the model the source so it doesn't have to guess.
An assistant suggests calling client.fetchAllPages() from a library you use. The name looks perfectly reasonable, but your build fails because the method doesn't exist. What happened?
Trained once, then it just runs
It helps to separate two very different phases. Training is the slow, expensive process where the model's internal settings are tuned by exposure to huge amounts of text — this happens once, ahead of time, in a data center. Inference is what happens when you actually use it: the finished model takes your input and predicts tokens. The key thing to internalize is that inference does not change the model. Asking it something today teaches it nothing for tomorrow; the model that answers your next question is byte-for-byte the same one.
No memory between calls
This is the consequence that trips people up most. Because using the model doesn't change it, a model has no memory of past conversations. Each call starts cold. If it seems to "remember" what you said three messages ago, that's only because all of those earlier messages are being re-sent to it every single time. Anything the model needs to know — the conversation so far, the contents of a file, your preferences — has to be packed into the input on each call. That input has a size limit, which is exactly what the context window lesson is about.