Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Context and Memory

How agents manage the context window, handle overflow, persist state, and spill large outputs.

Context Overflow Detection

Every agent monitors token usage against the context window:

AgentThresholdDetection
Goose80% of windowcheck_if_compaction_needed()
CodexConfigurableMultiple strategies selected at runtime
Grok BuildConfigurable per-agentshould_auto_compact(total_tokens, context_window, threshold)
Kimi CodeInfrastructure-levelTranscript ops handle overflow
OpenCode/KilocodeService-basedSessionCompaction effect

Compaction Strategies

Codex — Minimal Handoff

Prompt (~10 lines): “Create a handoff summary for another LLM that will resume the task.”

  • Include: progress, decisions, constraints, next steps, critical data
  • Output: free-form text
  • Multiple implementations: compact.rs, compact_remote.rs, compact_remote_v2_attempt.rs

Goose — Structured JSON

Prompt (~45 lines): Detailed section-by-section requirements.

  • Wrap reasoning in <analysis> tags (discarded)
  • Output JSON with 7 fields: user_intent, technical_concepts, files, errors_and_fixes, problem_solving, user_messages, pending_tasks
  • Rules: order by importance, quote errors verbatim, no new ideas
  • Continuation markers differ by context: tool-loop vs conversation vs manual

Grok Build — 9-Section Summary

Prompt (~20 lines): Numbered sections inside <summary> XML.

  1. Primary Request and Intent
  2. Key Technical Concepts
  3. Tool Usage & Verification
  4. Files, Attachments, Images, Render Results & Code Artifacts
  5. Errors and Fixes
  6. Problem Solving
  7. All User Messages
  8. Pending Tasks
  9. Optional Next Step

Special handling: chained compactions carry forward from prior summaries. References /tmp/compaction/segment_*.md as out-of-band memory.

Pi — File-Aware Compaction

Tracks which files were read/modified across compaction boundaries:

interface CompactionDetails {
  readFiles: string[];
  modifiedFiles: string[];
}

Previous compaction’s file lists are carried forward into the new compaction.

Kimi Code — Transcript Infrastructure

Not LLM-summarization at all. Uses a multi-level transcript system:

  • L1: Agent-granular store
  • L2: Idempotent operations
  • L3: off/turn/block/delta subscription granularity
  • L4: Framework-free view registry
  • Cold rebuild from wire.jsonl as single source of truth
  • Op-batch sequencing with point-to-point catch-up

Session Persistence

AgentFormatStructure
CodexJSONL (append-only)session-{id}.jsonl — each line is an event
GooseSQLite v15sessions.db with schema migrations
Grok BuildSQLite (WAL)Per-session with journal mode selection
Kimi CodeJSONL (wire.jsonl)Per-agent file, append + rewrite for compaction
PiJSONL (tree)parentId/leafId structure for branching
OpenCode/KilocodeSQLite (Drizzle)Effect-TS managed database layer
Qwen CodeJSONL + .outputSession dir with separate output files

JSONL vs SQLite

JSONL (Codex, Kimi Code, Pi, Qwen Code):

  • Append-only log — simple, streamable, easy to replay
  • Pi adds tree structure for branching (parent/leaf pointers)
  • Kimi Code supports rewrite for compaction

SQLite (Goose, Grok Build, OpenCode):

  • Schema migrations (Goose at v15)
  • Concurrent access safe
  • Structured queries for session listing/search
  • Grok Build selects journal mode (WAL vs rollback) based on filesystem type

Large Output Spilling

When tool output exceeds context budget, spill to filesystem:

AgentThresholdHead/TailFile Pattern
Codex~2,500 tokenshead + tail preview<temp>/hook_outputs/<thread_id>/<uuid>
Goose200,000 charsNo split — just pathgoose_mcp_response_*.txt
Qwen Code30,000 chars1/5 head + 4/5 tail<projectTemp>/<tool>.output

Universal pattern:

if output.size > threshold:
    path = write_to_temp(full_output)
    model_sees = truncated_preview + "Full output at: {path}"

The model can then use the read tool to examine the file — creating a feedback loop where the filesystem is working memory.

Background Task Output

Long-running commands bridge to the agent via filesystem:

  • Qwen Code: shell-${shellId}.output — stdout streamed to file, model reads via task_output
  • Codex: Hook outputs persisted per-thread under temp dir
  • Goose: Scheduled recipe executions persist as full sessions

The Memory Hierarchy

┌─────────────────────────────────────────────────┐
│  L1: Context Window (fast, limited, volatile)    │
├─────────────────────────────────────────────────┤
│  L2: Spilled Output Files (temp, read-on-demand) │
├─────────────────────────────────────────────────┤
│  L3: Session Persistence (durable, replayable)   │
├─────────────────────────────────────────────────┤
│  L4: Project Files (the actual codebase)         │
└─────────────────────────────────────────────────┘
  • L1 → L2: Overflow (large output spill)
  • L1 → L3: Checkpoint (compaction/session save)
  • L2 → L1: Recovery (read tool pulls spilled content back)
  • L4 → L1: Normal code reading flow

Token Optimization Strategies

  • Codex: Prompt cache prewarm — pre-caches system prompt before user types
  • Goose: Tool-pair summarization — batches of 10 old tool call/results compressed
  • Grok Build: xai-token-estimation crate for accurate counting; circuit breaker for API failures
  • Kimi Code: Subscription granularity (don’t send data the client won’t render)
  • Qwen Code: Microcompaction service for incremental context trimming