Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Universal Architecture — What Every Agent Has In Common

Every coding agent, regardless of language (Rust/TS/Python), interface (CLI/TUI/IDE/Web), or LLM provider, shares the same fundamental architecture. The differences are in how they implement these pieces, not in whether they have them.

The Core Loop (present in ALL agents)

┌─────────────────────────────────────────────────────────┐
│                     AGENT LOOP                          │
│                                                         │
│  1. Assemble context (system prompt + messages + tools) │
│  2. Call LLM (stream response)                          │
│  3. Parse response → text OR tool_calls                 │
│  4. If tool_calls: execute tools, append results, GOTO 2│
│  5. If text-only (end_turn): yield to user              │
│                                                         │
│  Guards:                                                │
│    - Max turns/steps limit                              │
│    - Abort/cancel signal                                │
│    - Context overflow → compact → retry                 │
│    - Error → retry with backoff                         │
└─────────────────────────────────────────────────────────┘

Evidence across codebases:

  • Pi (agent-loop.ts:169): outer while(true) with inner tool-call processing loop
  • Goose (agents/agent.rs:1948): loop { stream_response → process tool calls → check exit conditions }
  • Kimi Code (loop/run-turn.ts:136): while (true) { signal.throwIfAborted(); executeLoopStep(); }
  • Codex (codex_thread.rs): turn-based with session loop
  • Grok Build (acp_session.rs): run_turn_via_sampler with streaming turn capture

The 7 Universal Components

1. System Prompt Assembly

Every agent builds a system prompt from:

  • Base identity/personality (static)
  • Environment info (working dir, platform, date)
  • Project instructions (AGENTS.md / .claude / .goosehints)
  • Available tools description
  • Dynamic context (skill guidance, memory, prior compaction)

2. Message History (Conversation/Transcript)

Every agent maintains an ordered list of messages:

  • user → assistant → tool_result → assistant → …
  • Messages carry: role, content (text/image/tool_call/tool_result), metadata (usage, timing)
  • The history IS the context window content

3. Tool Registry + Execution

Every agent has a tool abstraction with the same shape:

Tool {
  name: string
  description: string
  parameters: JSONSchema
  execute(params) → result
}

Pi: AgentTool<TParameters> with execute(), label, executionMode Goose: Tool (from rmcp) with ToolAnnotations Grok Build: ToolDefinition via ToolBridge Kimi Code: ExecutableTool in the loop layer

Tool execution is either sequential or parallel. All agents track which tools are “in-flight.”

4. Context Overflow / Compaction

Every agent must handle “context too long” and all use the same basic strategy:

  • Detect: token count approaching context window (typically 80% threshold)
  • Summarize: call LLM to compress conversation history into a summary
  • Replace: swap history with summary + continuation marker
  • Resume: next turn sees summary instead of full history

5. Streaming + Events

Every agent streams LLM output token-by-token and emits typed events:

  • turn_start / turn_end
  • message_start / text chunks / message_end
  • tool_call / tool_result
  • usage (token counts)
  • error

6. Abort/Cancel Signal

Every agent has a cooperative cancellation mechanism:

  • Pi: AbortSignal checked between steps
  • Goose: CancellationToken checked in loop
  • Kimi Code: signal.throwIfAborted() at loop boundary
  • Grok Build: token-based cancellation

7. Session Persistence

Every agent persists conversation state across turns:

  • Codex: thread/session JSONL files
  • Goose: SessionManager with SQLite
  • Kimi Code: wire.jsonl records
  • Grok Build: SQLite journal
  • Pi: session JSONL entry files
  • OpenCode/Kilocode: Effect-TS database layer (Drizzle + SQLite)

Universal System Prompt Patterns

Despite different wording, every agent’s system prompt contains these same sections:

SectionWhat it saysPresent in
Identity“You are X, a coding assistant”All
Autonomy“Keep going until done”All
Tool preference“Use specialized tools over bash”All except Goose (delegated)
File editingHow to edit files (patch/edit/replace)All
Output styleBe concise, use markdownAll
Project instructionsRead and obey AGENTS.md / project filesAll
VerificationRun tests after changesAll
SafetyDon’t break things, confirm risky actionsAll

Universal Tool Categories

Despite different names, every agent provides tools in these categories:

CategoryPurposeExamples
File ReadRead file contentsread, read_file, cat
File Write/EditModify filesedit, write, apply_patch, search_replace
ShellExecute commandsbash, shell, run_commands
SearchFind in codebasegrep, rg, glob, search_codebase
DirectoryList filesls, list_dir, glob

These 5 categories are present in every agent. Everything else (web, plan, memory, subagent, workflow) is bonus.

File System as Architecture

File reading and writing aren’t just “tools” — they define the agent’s fundamental relationship with code. The read/write strategy determines:

  • What the LLM sees (line numbers? anchors? diffs? raw content?)
  • How edits are specified (exact match? line numbers? anchors? whole-file?)
  • What can go wrong (stale references, ambiguous matches, merge conflicts)

The Read→Edit Coupling (present in ALL agents)

Every agent couples its read format to its edit format. The read tool produces output that the edit tool consumes as addressing:

AgentRead FormatEdit AddressingEdit Mechanism
Codexraw (via shell)Context lines (3 before/after)apply_patch — custom diff (@@-anchored hunks, +/- lines)
ClineLINE_NUMBER→CONTENTexact old_text match OR insert_lineedit_file (search/replace) + apply_patch (diff grammar)
Goose(via MCP extensions)(extension-dependent)(extension-dependent)
Grok BuildLINE:HASH:CTX→CONTENTAnchor-based (22:abc:rst)Hashline edit (atomic batch, stale = reject all)
Grok Build (alt)LINE_NUMBER→CONTENTexact old_string matchsearch_replace (find & replace)
Kimi Code(dynamic)(dynamic)(dynamic)
OpenCode/Kilocoderaw with line numbersexact oldString matchedit (search/replace with replaceAll option)
Piraw with line numbersexact old_text matchedit tool
Qwen Codecat -n format (line numbers)exact old_string matchedit (search/replace)

Three Families of Edit Strategy

1. Exact String Match (most common)

  • Used by: Cline, OpenCode, Kilocode, Qwen Code, Pi, Grok Build (search_replace)
  • old_string must match exactly once in the file → replaced with new_string
  • old_string = "" or null → create new file
  • Failure mode: ambiguous match (appears 0 or >1 times)
  • Mitigation: replaceAll flag, or “add surrounding lines to make unique”

2. Diff/Patch Language

  • Used by: Codex, Cline (apply_patch)
  • Custom mini-language with *** Begin Patch, @@ context, +/- lines
  • Addresses via context (like git diff) not exact match
  • Failure mode: context doesn’t match current file state
  • Advantage: can express multi-hunk changes in a single tool call

3. Anchor-Based (unique to Grok Build)

  • Each line gets a content-derived hash anchor: LINE:HASH→CONTENT
  • Edits reference anchors, not line numbers or string matches
  • Atomic batch semantics: if any anchor is stale, ALL edits rejected
  • Stale anchor → error response includes fresh anchors → model retries immediately
  • Three scheme candidates: ContentOnly, ChunkFingerprint, CheckpointChain
  • Advantage: robust to concurrent edits / line shifts
  • Disadvantage: anchor churn after edits, more complex protocol

The File Creation Pattern

Every agent needs to handle “create new file” distinctly from “edit existing file”:

  • Codex: *** Add File: <path> header in patch language
  • Cline/OpenCode/Qwen: old_string = null/empty + new_string = content
  • Grok Build (search_replace): old_string = "" creates file
  • Pi: separate write tool for full-file writes

Why This Is Architectural

The read/edit coupling shapes the entire agent experience:

  1. Token efficiency: Codex sends minimal context (3 lines); hashline requires anchors for every line read
  2. Reliability: exact-match can fail on repeated code; anchors handle line shifts gracefully
  3. Multi-edit atomicity: patch language batches hunks; search_replace is one-at-a-time; hashline batches are atomic
  4. Model burden: exact-match requires the model to reproduce code perfectly; patch format allows context-based targeting
  5. Read-before-write requirement: Nearly all agents enforce “must read before edit” (OpenCode errors if not, Grok Build’s prompt says “Read the file first”)

The Filesystem as Infrastructure

The filesystem isn’t just a target of code edits — it’s a core piece of the agent’s own infrastructure. Every agent uses files for three internal purposes beyond code editing:

A. Session Persistence (conversation state across turns)

Every agent persists conversation history to disk so sessions survive restarts:

AgentStorage FormatLocation
CodexJSONL (append-only)session-{id}.jsonl
GooseSQLite (v15 schema)~/.config/goose/sessions.db
Grok BuildSQLite (WAL/journal)Session directory
Kimi CodeJSONL (wire.jsonl)Per-agent file, append + rewrite
PiJSONL (append-only tree)Per-session .jsonl file
OpenCode/KilocodeSQLite (via Drizzle + Effect)Database layer
Qwen CodeJSONL + output filesSession directory

The two strategies:

  • JSONL (Codex, Kimi Code, Pi, Qwen Code): Append-only log of records. Simple, streamable, easy to replay. Pi adds tree structure (parentId) for branching.
  • SQLite (Goose, Grok Build, OpenCode): Structured queries, schema migrations, concurrent access. Goose is at schema v15 — it evolves.

B. Large Output Spilling (keeping context manageable)

When a tool produces output too large for the context window, every agent writes it to a temp file and replaces it with a pointer:

AgentThresholdStrategyFile Location
Codex~2,500 tokensSpill to file, show head/tail preview + path<temp>/hook_outputs/<thread_id>/
Goose200,000 charsWrite full output to file, replace with file path messagegoose_mcp_response_*.txt (tempfile)
Qwen Code30,000 chars (shell)head 1/5 + tail 4/5, save full to .output file<projectTemp>/<toolName>.output
Grok Build(configurable)Per-tool truncation with output filesSession temp directory

The universal pattern:

if output.size > threshold:
    file = write_to_temp(full_output)
    model_sees = f"Output too large ({size}). Saved to: {file}\n{head}...[TRUNCATED]...{tail}"

This creates a feedback loop with file reading: the model can then use the read tool to examine the spilled file if it needs the full content. The filesystem becomes a working memory extension.

C. Compaction Persistence (surviving context resets)

Some agents persist compaction artifacts beyond the current context:

  • Grok Build: References /tmp/compaction/segment_*.md and /tmp/compaction/INDEX.md as “out-of-band memory channels for a future work agent.”
  • Kimi Code: Plan versions persisted as agents/<agentId>/plan/<planId>/v<N>.md with SHA256 content hash. Cold rebuild from wire.jsonl.
  • Codex: Session logs enable resume from any point.
  • Pi: Compaction entries stored in the session JSONL with file operation tracking (which files were read/modified).

D. Background Task Output

When agents run long-running commands in the background, the filesystem bridges the gap:

  • Qwen Code (shell.ts:3058): shell-${shellId}.output — background shell stdout streamed to file, model reads later via task_output tool
  • Codex: Hook outputs spilled to disk per-thread
  • Goose: Scheduled recipe execution results in sessions

Why This Matters Architecturally

The filesystem serves as the agent’s external memory hierarchy:

┌─────────────────────────────────────────────────┐
│  L1: Context Window (fast, limited, volatile)    │
├─────────────────────────────────────────────────┤
│  L2: Spilled Output Files (temp, read-on-demand) │
├─────────────────────────────────────────────────┤
│  L3: Session Persistence (durable, replayable)   │
├─────────────────────────────────────────────────┤
│  L4: Project Files (the actual codebase)         │
└─────────────────────────────────────────────────┘
  • L1 → L2: Overflow. When tool output exceeds context budget, spill to file. Model can read back if needed.
  • L1 → L3: Checkpoint. On compaction or session save, persist full state so it can be restored.
  • L2 → L1: Recovery. Model uses read tool to pull spilled content back into context when needed.
  • L4 → L1: The normal read-file-into-context flow for code editing.

This is functionally the same architecture as CPU cache hierarchies — the context window IS the L1 cache, and the filesystem is everything slower but larger.

The Conversation Shape

Every agent uses the same message shape for LLM communication:

[
  { role: "system", content: [assembled system prompt] },
  { role: "user", content: "user's request" },
  { role: "assistant", content: "thinking + tool_calls" },
  { role: "tool", content: "tool results" },
  { role: "assistant", content: "more tool_calls or final response" },
  ...
]

The only variation is whether thinking/reasoning is a separate content block or inline.

Key Structural Insight

The entire architecture can be reduced to:

An agentic coding CLI is a loop that repeatedly calls an LLM with (system_prompt + history + tool_definitions), executes any tool_calls in the response, appends results to history, and repeats until the LLM produces a text-only response or a guard fires.

Everything else — compaction, subagents, plan mode, skills, permissions, memory — is optimization or UX layered on top of this core loop.