Context window management
Every model has a fixed context window: the maximum number of tokens it can read in a single request. A long conversation, or one with many tool calls and large tool results, accumulates tokens with every turn and can eventually exceed that window. Context window management keeps a conversation within the window so it continues without interruption.
Automatic compaction
When a conversation approaches the context window, the gateway compacts it automatically. Compaction summarises and condenses the earlier part of the history so the most recent turns are preserved in full, while the older content is reduced to fit the remaining budget. On the /easy surface the conversation simply continues; there is no manual step.
On the chat (/easy) surface the summarize pass is incremental and non-blocking: the summary is streamed in the background as it is produced — a brief Compacting earlier conversation… indicator shows while it runs, and Context compacted — summary saved when it finishes — and the user can keep reading and typing throughout. The condensed history is persisted as it forms, so a long-running conversation does not stall on a single large summarisation step.
There are two compaction mechanisms. On the chat (/easy) surface, a model-agnostic summarize pass condenses older turns as the conversation grows — this works for any model. Separately, for Anthropic models that support Anthropic's native compact_20260112 strategy, the gateway also requests native compaction on the upstream call and the provider decides what to keep and what to summarise. Anthropic models that do not support that strategy (for example claude-haiku-4-5, claude-sonnet-4-5, claude-opus-4-5) rely on the summarize pass on /easy; on a raw API call such a model is not compacted by the native mechanism.
Rules and facts
- Compaction is enabled per gateway. When enabled, it triggers once the estimated input token count reaches the configured threshold.
- The default threshold is 200000 tokens. The Anthropic API requires a minimum of 50000 tokens, so a lower threshold has no effect below that floor.
- Native compaction applies to Anthropic models that support the
compact_20260112strategy. Other providers, and Anthropic models without that strategy, are not compacted by this mechanism (the model-agnostic/easysummarize pass still applies there).
Tool-loop safety limits
A single turn can run several rounds of tool calls. To stop a runaway tool loop, the gateway caps each turn by wall-clock time, by the number of rounds, and by repeated identical tool calls, so a turn can never hang.
When the wall-clock limit is reached after the model has already gathered some tool results, the turn does not simply stop with an error. Instead the gateway runs one final, bounded synthesis step that composes the best answer it can from the results collected so far, and returns it marked as a partial result. The user gets a useful, honest answer for the work that completed rather than nothing.
Long and data-file turns
Two further mechanisms keep turns that carry a spreadsheet or another large data file within the window without the user having to manage anything.
- Data-file preview bounding. When a conversation re-sends a large attached file, the inline preview of that file is bounded at request time so a re-sent preview cannot, by itself, overflow the window. The most useful head of the file is kept; the rest is trimmed for the purpose of the running context. The original file remains available to tools such as the code interpreter for exact computation.
- Automatic model upgrade for data-file turns. When a turn works over a data file and the current model is not provably suited to it — for example it lacks the tools or the context window to handle the file — the gateway upgrades that single turn to a capable model (one with the right tools and a larger window) so the file can actually be processed. The upgrade applies to that turn only; the rest of the conversation continues on the chosen model. On the
/easysurface this happens automatically under the Auto model setting.
See also
- Configuration reference — Context compaction — the
context_compactiongateway settings. - Anthropic — the provider that performs native compaction.