Imagine trying to answer a complex question while only being allowed to look at one page of notes at a time. Whatever isn't on that page may as well not exist. A language model lives under exactly this constraint. As we saw with large language models, a model has no memory between calls — so the text you hand it in a single request is, quite literally, everything it can think about. The container that holds that text has a fixed size, and that container is the context window.
A bounded budget of tokens
The window is measured in tokens — the same small text chunks the model reads and writes. Every model has a maximum: it can take in only so many tokens per call, full stop. Think of it less like a notebook you can keep adding pages to and more like a single whiteboard of a fixed size. You can write whatever you like on it, but once it's full, something has to be erased before anything new fits. Window sizes vary a lot between models, from tens of thousands of tokens to a million or more, but every one of them has a ceiling, and that hard ceiling shapes almost everything about working with these models.
How it works: what shares the space
A lot has to fit on that one whiteboard at the same time. There's the system prompt — the standing instructions that set how the model behaves. There's the conversation history — every previous turn, re-sent so the model appears to remember. There are the tool definitions that describe what actions are available. And there's whatever you've pasted in — files, error logs, documentation. All of it competes for the same fixed budget.
Here's the part that surprises people: because the model has no memory between calls, the whole conversation is re-sent on every turn. A session doesn't fill the window once; it refills it on every call, a little fuller each time. Step through a long session below. The numbers are deliberately small (a 40K-token window) so you can do the arithmetic. When the next turn won't fit, predict what the model loses before you look, then flip where the rule lives.
In our stack — when Claude Code works on your project, its context window is constantly being assembled: the instructions that guide it, the conversation so far, the definitions of the tools and any MCP connections it can use, and the slices of your codebase relevant to the task. The Claude model behind it has a generous window, but it is still finite — so Claude Code is deliberate about which files it pulls in rather than dumping the whole repository onto the whiteboard.
When the window fills up
Long sessions inevitably bump against the ceiling. When the total — instructions plus history plus tools plus files — would exceed the limit, something has to give, and there are two usual responses. The simplest is to drop the oldest turns, as in the scene above. The other is to summarize: compress earlier parts of the conversation into a short recap that keeps the gist and frees up tokens. Summaries are gentler, but they're still lossy. The recap of a two-hour session might keep "we're refactoring billing" and lose "and don't touch /legacy".
Either way, detail disappears, and it's usually the oldest detail: the instructions you gave at the very start, which are often the most important ones. This is why a model can seem to "forget" something you mentioned long ago in a marathon session. It didn't forget so much as the note got erased from the whiteboard to make room. Even before the hard limit, a single sentence buried under tens of thousands of tokens of later chat can carry less weight than you'd expect.
Rules said once, early, are the first to go. A constraint you mention in your first message lives in the oldest turn, which is exactly what gets dropped or summarized away when the window fills. Put rules that must always hold into the pinned instructions (a system prompt or project instructions file), which are re-sent with every call. For a one-off rule, restate it close to the request it applies to.
Two hours into a session, an assistant starts editing a folder you told it to leave alone in your very first message. What's the most likely cause?
Why this is worth understanding
Once you picture the window as a finite, shared space, a lot of behavior stops being mysterious and becomes manageable. Putting the most important instructions where they won't get crowded out, keeping irrelevant files out of the budget, and starting fresh when a conversation has grown bloated are all ways of respecting the limit. It also explains cost and speed: everything in the window is sent again on every call, so a bloated session tends to be slower and more expensive as well as more forgetful. Doing this deliberately — choosing what deserves a spot on the whiteboard and what doesn't — is its own discipline, which the context engineering lesson explores in depth.
Your context window is nearly full. Which of these is NOT taking up any of it?