Glossary · Agent architecture

Context window

The amount of text, measured in tokens, a language model can take into account at once, including the prompt, history, tool data and its own output.

The context window is the total amount of text, measured in tokens, that a language model can take into account when generating a response, including the response itself.

What counts. Anthropic’s documentation lists everything in a request: the system prompt, every message (including tool results, images and documents), the tool definitions, and the output the model generates, including its thinking. Cached tokens still count toward the window, even though prompt caching lowers what they cost. Window sizes differ by model and change often, so check the provider’s current model table.

Size and quality. A larger window lets a model handle longer inputs, but the same documentation notes that accuracy and recall degrade as the token count grows, which Anthropic calls context rot. What goes into the window matters as much as how large it is.

At the limit. Behaviour depends on the API. Anthropic’s API returns an error when the input alone exceeds the window, and on newer models stops generation with a dedicated stop reason if the output runs into the limit. Chat products may instead drop the oldest turns as the conversation grows.

Managing it in agents. Agents fill the window quickly, because every tool call adds a request and a result. Anthropic’s guidance on context engineering describes the main techniques:

  • Compaction: summarize earlier parts of the conversation and continue from the summary.
  • Structured note-taking: keep notes outside the window and reload them when relevant.
  • Sub-agents: do detailed work in a separate window and return a short summary.
  • Just-in-time retrieval: load data through tools when it is needed, using lightweight references until then.

Across agents. Between organizations, each agent has its own window. In A2A, agents exchange messages and artifacts, so the calling agent decides what to send, and the remote agent’s internal work never enters the caller’s window. Only what comes back does, and that content should be treated as untrusted input.

Neighbouring terms. Agent memory decides what gets loaded back into the window. Sub-agents are a way to spread work across several windows.

Sources

  1. Anthropic documentation: Context windows (accessed )
  2. Anthropic: Effective context engineering for AI agents (29 September 2025) (accessed )
  3. A2A Protocol Specification (accessed )