Context and Memory
How agents manage the context window, handle overflow, persist state, and spill large outputs.
Context Overflow Detection
Every agent monitors token usage against the context window:
| Agent | Threshold | Detection |
|---|---|---|
| Goose | 80% of window | check_if_compaction_needed() |
| Codex | Configurable | Multiple strategies selected at runtime |
| Grok Build | Configurable per-agent | should_auto_compact(total_tokens, context_window, threshold) |
| Kimi Code | Infrastructure-level | Transcript ops handle overflow |
| OpenCode/Kilocode | Service-based | SessionCompaction effect |
Compaction Strategies
Codex — Minimal Handoff
Prompt (~10 lines): “Create a handoff summary for another LLM that will resume the task.”
- Include: progress, decisions, constraints, next steps, critical data
- Output: free-form text
- Multiple implementations:
compact.rs,compact_remote.rs,compact_remote_v2_attempt.rs
Goose — Structured JSON
Prompt (~45 lines): Detailed section-by-section requirements.
- Wrap reasoning in
<analysis>tags (discarded) - Output JSON with 7 fields:
user_intent,technical_concepts,files,errors_and_fixes,problem_solving,user_messages,pending_tasks - Rules: order by importance, quote errors verbatim, no new ideas
- Continuation markers differ by context: tool-loop vs conversation vs manual
Grok Build — 9-Section Summary
Prompt (~20 lines): Numbered sections inside <summary> XML.
- Primary Request and Intent
- Key Technical Concepts
- Tool Usage & Verification
- Files, Attachments, Images, Render Results & Code Artifacts
- Errors and Fixes
- Problem Solving
- All User Messages
- Pending Tasks
- Optional Next Step
Special handling: chained compactions carry forward from prior summaries. References /tmp/compaction/segment_*.md as out-of-band memory.
Pi — File-Aware Compaction
Tracks which files were read/modified across compaction boundaries:
interface CompactionDetails {
readFiles: string[];
modifiedFiles: string[];
}
Previous compaction’s file lists are carried forward into the new compaction.
Kimi Code — Transcript Infrastructure
Not LLM-summarization at all. Uses a multi-level transcript system:
- L1: Agent-granular store
- L2: Idempotent operations
- L3:
off/turn/block/deltasubscription granularity - L4: Framework-free view registry
- Cold rebuild from
wire.jsonlas single source of truth - Op-batch sequencing with point-to-point catch-up
Session Persistence
| Agent | Format | Structure |
|---|---|---|
| Codex | JSONL (append-only) | session-{id}.jsonl — each line is an event |
| Goose | SQLite v15 | sessions.db with schema migrations |
| Grok Build | SQLite (WAL) | Per-session with journal mode selection |
| Kimi Code | JSONL (wire.jsonl) | Per-agent file, append + rewrite for compaction |
| Pi | JSONL (tree) | parentId/leafId structure for branching |
| OpenCode/Kilocode | SQLite (Drizzle) | Effect-TS managed database layer |
| Qwen Code | JSONL + .output | Session dir with separate output files |
JSONL vs SQLite
JSONL (Codex, Kimi Code, Pi, Qwen Code):
- Append-only log — simple, streamable, easy to replay
- Pi adds tree structure for branching (parent/leaf pointers)
- Kimi Code supports rewrite for compaction
SQLite (Goose, Grok Build, OpenCode):
- Schema migrations (Goose at v15)
- Concurrent access safe
- Structured queries for session listing/search
- Grok Build selects journal mode (WAL vs rollback) based on filesystem type
Large Output Spilling
When tool output exceeds context budget, spill to filesystem:
| Agent | Threshold | Head/Tail | File Pattern |
|---|---|---|---|
| Codex | ~2,500 tokens | head + tail preview | <temp>/hook_outputs/<thread_id>/<uuid> |
| Goose | 200,000 chars | No split — just path | goose_mcp_response_*.txt |
| Qwen Code | 30,000 chars | 1/5 head + 4/5 tail | <projectTemp>/<tool>.output |
Universal pattern:
if output.size > threshold:
path = write_to_temp(full_output)
model_sees = truncated_preview + "Full output at: {path}"
The model can then use the read tool to examine the file — creating a feedback loop where the filesystem is working memory.
Background Task Output
Long-running commands bridge to the agent via filesystem:
- Qwen Code:
shell-${shellId}.output— stdout streamed to file, model reads viatask_output - Codex: Hook outputs persisted per-thread under temp dir
- Goose: Scheduled recipe executions persist as full sessions
The Memory Hierarchy
┌─────────────────────────────────────────────────┐
│ L1: Context Window (fast, limited, volatile) │
├─────────────────────────────────────────────────┤
│ L2: Spilled Output Files (temp, read-on-demand) │
├─────────────────────────────────────────────────┤
│ L3: Session Persistence (durable, replayable) │
├─────────────────────────────────────────────────┤
│ L4: Project Files (the actual codebase) │
└─────────────────────────────────────────────────┘
- L1 → L2: Overflow (large output spill)
- L1 → L3: Checkpoint (compaction/session save)
- L2 → L1: Recovery (read tool pulls spilled content back)
- L4 → L1: Normal code reading flow
Token Optimization Strategies
- Codex: Prompt cache prewarm — pre-caches system prompt before user types
- Goose: Tool-pair summarization — batches of 10 old tool call/results compressed
- Grok Build:
xai-token-estimationcrate for accurate counting; circuit breaker for API failures - Kimi Code: Subscription granularity (don’t send data the client won’t render)
- Qwen Code: Microcompaction service for incremental context trimming