What is a context window? (and why AI forgets)
If you've used AI for anything longer than a quick question, you've probably hit this: halfway through a long back-and-forth it starts contradicting something you established at the beginning, or asks for a detail you already gave it. It looks like carelessness. It's actually a hard limit doing exactly what it does.
Think of it as a desk, not a filing cabinet
A model doesn't have memory in the way a person does. It has a working surface — a desk — and everything it can consider has to be on that desk at once. Your instructions, the whole conversation, the document you attached, and every reply it has already given.
The desk is large but finite. When you add something new and there's no room, the material at the far edge slides off. It isn't filed away for later. It's simply gone from view.
This is why "remember that our financial year ends in June" works reliably in a short exchange and unreliably in a two-hour one.
What counts toward the limit
Everything in the exchange, not just what you type:
- Your prompts
- The model's own replies — often the biggest consumer, since they're usually longer than your questions
- Attached files and pasted material, in full
- Any standing instructions the product adds behind the scenes
A single 40-page PDF can occupy a serious share of the window on its own. Two of them, plus an hour of discussion, and you are closer to the edge than you'd guess.
An Australian small-business example
A Brisbane bookkeeper starts a session by explaining the client's chart of accounts, their GST treatment, and three unusual rules that apply to this client only. The first twenty categorisations are perfect.
By the sixtieth, the model starts applying the standard treatment instead of the client-specific rule. Nothing broke — the explanation from the start of the session is no longer on the desk, crowded out by sixty transactions and sixty replies.
The fix isn't a better prompt. It's a shorter session: handle twenty transactions, start fresh, restate the three rules. Same work, a fraction of the errors.
Working with the limit instead of against it
- Put the important constraints last. Material at the very start and the very end of a window carries the most weight; the middle is where things get lost. If something must not be missed, say it close to the request.
- Start new sessions at natural breaks. A fresh window with a three-line brief beats a long thread carrying an hour of noise.
- Attach the relevant pages, not the whole document. Ten pages that matter will beat two hundred that mostly don't — for accuracy, not just speed.
- Re-state rather than assume. If a rule is critical and you're deep in a conversation, repeat it. It costs one sentence.
- Ask it to summarise before you switch topics. Then carry that summary into the next session. You're compressing the context deliberately instead of letting it fall off the edge.
Why this matters more than it sounds
Most complaints that an AI tool is "unreliable" turn out to be context problems rather than capability problems. The model didn't get worse — it stopped being able to see the thing that made the answer right.
Once a team understands the desk metaphor, their results improve without any change in tool, subscription or skill. They just stop asking a model to remember something it was never holding.
Frequently asked questions
How big is a context window?
Why did the AI forget what I told it ten minutes ago?
Does a bigger context window always mean better answers?
Does the AI remember between conversations?
What's the practical fix for long projects?
Put this to work
Ad On Group runs AI training and enablement for Australian teams through Ad On AI — a three-month, self-paced program that takes non-technical staff from their first prompts to working AI agents.