Context Window
The context window states how much text a model can take into account at once — instruction, documents, conversation history and answer combined. It is measured in tokens.
How it works
The context window is the model’s working memory. Everything it should consider for an answer — your instruction, uploaded documents, the conversation so far and the answer itself — has to fit in. Its size is given in tokens and, depending on the model, ranges from a few thousand to hundreds of thousands of tokens and more.
A practical example
You want to review a long contract with a model. If it fits entirely into the context window, the model can read it as a whole. If it is longer, you have to split it up or pass only the relevant passages via RAG.
What you should know
- A bigger window does not automatically mean better answers: with very long inputs, details are more easily overlooked.
- More context costs more tokens and therefore more money and time.
- What lies outside the window is unknown to the model — even from an earlier conversation, unless it is stored and passed in again.