Skip to main content
Every model has a fixed context window — the amount of conversation it can “see” at once. A long agent session eventually fills it. The naive fix is to summarize the whole history into a paragraph the moment it gets tight — one big, lossy step that throws away detail and often drops the thread. Monad instead relieves the pressure in a cascade: the cheapest, least destructive thing first, and only escalating when it must — so a session can run for hours without a jarring “I forgot everything” moment. Two optional stages sit alongside the cascade: a recitation anchor that re-pins the current plan after a summary, and semantic retrieval that pulls relevant older messages back when a new question needs them. And memory promotion can lift durable facts out of a span just before it’s compacted away.

For operators — do I need to touch this?

No, by default. The lossless and lossy stages are on out of the box with sensible thresholds; a fresh install manages its own context and you never have to think about it. Everything lives under the context.* block in agents.json and hot-reloads on save (no restart). You’d reach for the config only to: The full Configuration table lists every knob. The rest of this document is for developers working on the cascade itself.

The pressure ladder (developers)

Each stage triggers at a fraction of the active model’s context limit. Occupancy is measured by effectiveInputTokens (context/index.ts): it prefers the provider’s real input-token count from the previous step (true occupancy — system prompt + tool schemas + messages) and falls back to a self-calibrating char/token estimate before any real count exists. The one-step lag is fine for soft/eviction triggers; the hard overflow guard stays on the in-turn estimate, which tracks growth without lag. Tool-output truncation (below) is separate — it fires per result at a fixed char cap, not on window pressure. Code map:

Two assembly phases

The stages don’t all run at the same moment. There are two, and getting a stage in the wrong phase either duplicates injected text or misses mid-loop growth.
A stage that appends to the prompt (retrieval, recitation) runs per-turn to avoid duplicating its block on every step; a stage that reacts to the growing window (eviction, token limiter) runs per-step. Regression tests pin this in test/unit/agent/retrieval.test.ts.

Stage 1 — lossless tool-result eviction

Tool results (file reads, command output, search dumps) dominate a long transcript and go stale fast: once the model has acted on a read_file, the raw bytes rarely matter again. ToolResultEvictionContext reclaims that space without summarizing by replacing the output of old tool-result parts with a placeholder ([context-cleared] …). The tool-call/tool-result pairing is preserved (only the output text changes, no message dropped), so strict providers never see an orphan.
  • Recency is measured in ROUNDS, not results. A round is one assistant→tools step; parallel tool calls in one step land in one tool message (same age). Protecting the last keepRecentRounds rounds keeps each recent step whole regardless of how many concurrent results it produced — a flat “keep N results” count could split one concurrent batch.
  • It fires in batches. Only once occupancy crosses atFraction and a single pass can reclaim ≥ clearAtLeast tokens; results smaller than minResultTokens are skipped (not worth a placeholder).
  • Idempotent. An already-evicted placeholder is never re-evicted.
The placeholder points the model at read_tool_output when a recovery handle exists (next section) — so eviction upgrades from “re-run the tool (maybe non-reproducible or side-effecting)” to “read the exact bytes back.”

Recovery by handle — tool_raw_outputs + read_tool_output

Truncation and eviction both hide bytes the model might still need. Both spill the full pre-truncation output to a store so it can be paged back deterministically instead of re-run. Tool-output truncation (at execution time). truncateToolOutput (loop/tool-output.ts) caps a single result at toolOutput.maxChars (24k default) with a head(70%)+tail(30%) window and an embedded read_tool_output({ id, offset, limit }) hint. When toolOutput.persistRaw is on, the full output (capped at rawCapBytes) is written to tool_raw_outputs keyed by (transcript_target_id, tool_call_id). The read_tool_output tool. Pages back a spilled output by its tool-call id, with offset/limit or a grep substring filter, re-truncated at the same cap so paging can’t itself blow the window. Registered only when persistRaw is on (nothing is ever spilled otherwise, so the tool would always report “not found”). Store semantics (store/db/index.ts):
  • saveToolRawOutput upserts on the primary key.
  • getToolRawOutput reads by exact (transcript_target_id, tool_call_id) — scoped to one transcript, so one session can’t read another’s bytes.
  • Branching copies, not shares. Session branching clones the message snapshot (cloneMessages); the branch handler then calls cloneToolRawOutputs to copy the snapshot tool calls’ spilled outputs into the child, so the cloned tool_call rows’ read_tool_output handles keep resolving against the child’s own id. The copies are independent — overwriting the child’s doesn’t touch the parent’s.
  • Cleaned up on session/project delete and on session reset (cascade in store/db/sessions.ts).
Eviction’s own spill covers results that were never truncated at execution time (short enough to send whole), so evicting an old short result still leaves it recoverable. It skips re-spilling a result that already has a raw output (checked via hasRawOutput) — by eviction time the in-prompt text is already the persisted copy, so re-spilling would overwrite the correct full-text entry with a truncated one.

Redaction safety

An AfterTool hook may redact secrets from a tool result. The hook runs once, against the full pre-truncation text — not a truncated preview — so a secret in the omitted middle of a large output still triggers redaction; the (possibly rewritten) result is truncated afterward. When the hook redacts, no raw output is spilled and no recovery handle is advertised, so a redacted secret can’t leak back through read_tool_output.

Stage 2 — durable summarization

DurableSummarizer (history.ts) is the lossy stage. Instead of loading the whole transcript and trimming, it loads only messages since a durable summary boundary plus a persisted rolling summary — so per-turn DB read and memory stay O(window) regardless of session length, and both survive restarts.
  • The summary is a structured briefing (`## Objective / ## Decisions & Facts /

    Files & State / ## Open Tasks / ## Next Step`), with identifiers quoted

    verbatim — not free prose.
  • keepRecent most-recent messages are always sent verbatim (never summarized).
  • Background mode (default). Crossing softFraction kicks a non-blocking compaction (deduped per session); the turn proceeds at full width and the durable result lands whenever it finishes. A turn at/over hardFraction waits for any in-flight compaction, then compacts synchronously if still over.
  • The summary is folded into the first user message (not a separate system message — splitSystem keeps only the first system message, so a second would be dropped). The cheapest configured model (the fast tier) does the summarization.
  • A reflect pass GC-condenses the rolling summary once it exceeds ~4000 tokens.
/compact forces a compaction immediately regardless of threshold.

Stage 3 — recitation anchor (opt-in)

After compaction, “what am I doing right now” is buried in dense summary prose. When recitation.enabled, parsePlanSections (context/recitation.ts) pulls the summary’s ## Open Tasks / ## Next Step sections and appends a <plan> anchor to the end of the prompt, closest to where the model generates. No-op without a summary or when both sections are absent.

Stage 4 — semantic retrieval reinjection (opt-in)

Eviction and summarization shrink the sent window, but the store always keeps every message’s original text. So a semantic search over the session’s full history can recover exactly what the lossy/lossless stages hid, when a later turn needs it again. When retrieval.enabled (and an embedding model is configured), RetrievalReinjectionContext (context/retrieval.ts) embeds the latest user message, searchSemantic-es this session’s history (scoped by transcript target — no cross-session leakage), and splices hits ≥ minScore (capped at maxResults, pre-truncated snippets) into a <related_context> block on the user turn. Runs once per turn (see two phases); no-ops silently when embedding is unavailable so a retrieval hiccup never fails a turn.

Memory promotion (opt-in)

A span about to be compacted away is the last chance to pull durable facts out of its original text (the summary is already a paraphrase). When memoryPromotion.mode is not off, DurableSummarizer.afterCompact hands the folded transcript to MemoryService.promoteFacts (services/memory/index.ts), a fast-model extraction of durable facts:
  • suggest — emits a persisted memory.suggestion event; the user confirms via the web toast, which writes through the normal addMemoryFact path.
  • auto — writes straight to the resolved scope (the session’s agent, or global when it has none).
Both modes go through the same sanitizeFact chokepoint as the manual memory tool (memory.md) — no weaker validation on the auto path. Best-effort: a failed extraction never affects the turn.

Telemetry and UX

Four wire events surface what the cascade did (all data-plane, over per-session SSE — see realtime-channels.md): context.evicted / context.handoff_suggested are filtered from persistence in handlers/session/context.ts alongside session.message.delta.appended — transient notices, not history. Web. The composer’s context-usage panel shows the reclaimed line when > 0. useContextNotices turns the handoff nudge into a toast and the memory suggestion into a “remember these N things?” prompt with a Save action. context.evicted is intentionally not toasted — it’s routine housekeeping already visible in the usage panel and would fire on nearly every turn once the window fills.

Configuration

All under context.*; hot-reloaded. Fractions are of the active model’s context limit.

Reverting to pre-cascade behavior

Each stage is independently disable-able:
  • eviction.enabled: false drops the lossless stage (summarization then carries the full load).
  • summarize.background: false makes compaction synchronous at the soft threshold.
  • toolOutput.persistRaw: false disables recovery-by-handle (truncation still caps the model-visible result; read_tool_output is not registered).
  • The opt-in stages (recitation, memory promotion, handoff nudge, retrieval) are off by default.