Universal Architecture — What Every Agent Has In Common
Every coding agent, regardless of language (Rust/TS/Python), interface (CLI/TUI/IDE/Web), or LLM provider, shares the same fundamental architecture. The differences are in how they implement these pieces, not in whether they have them.
The Core Loop (present in ALL agents)
┌─────────────────────────────────────────────────────────┐
│ AGENT LOOP │
│ │
│ 1. Assemble context (system prompt + messages + tools) │
│ 2. Call LLM (stream response) │
│ 3. Parse response → text OR tool_calls │
│ 4. If tool_calls: execute tools, append results, GOTO 2│
│ 5. If text-only (end_turn): yield to user │
│ │
│ Guards: │
│ - Max turns/steps limit │
│ - Abort/cancel signal │
│ - Context overflow → compact → retry │
│ - Error → retry with backoff │
└─────────────────────────────────────────────────────────┘
Evidence across codebases:
- Pi (
agent-loop.ts:169): outer while(true) with inner tool-call processing loop - Goose (
agents/agent.rs:1948):loop { stream_response → process tool calls → check exit conditions } - Kimi Code (
loop/run-turn.ts:136):while (true) { signal.throwIfAborted(); executeLoopStep(); } - Codex (
codex_thread.rs): turn-based with session loop - Grok Build (
acp_session.rs):run_turn_via_samplerwith streaming turn capture
The 7 Universal Components
1. System Prompt Assembly
Every agent builds a system prompt from:
- Base identity/personality (static)
- Environment info (working dir, platform, date)
- Project instructions (AGENTS.md / .claude / .goosehints)
- Available tools description
- Dynamic context (skill guidance, memory, prior compaction)
2. Message History (Conversation/Transcript)
Every agent maintains an ordered list of messages:
user→assistant→tool_result→assistant→ …- Messages carry: role, content (text/image/tool_call/tool_result), metadata (usage, timing)
- The history IS the context window content
3. Tool Registry + Execution
Every agent has a tool abstraction with the same shape:
Tool {
name: string
description: string
parameters: JSONSchema
execute(params) → result
}
Pi: AgentTool<TParameters> with execute(), label, executionMode
Goose: Tool (from rmcp) with ToolAnnotations
Grok Build: ToolDefinition via ToolBridge
Kimi Code: ExecutableTool in the loop layer
Tool execution is either sequential or parallel. All agents track which tools are “in-flight.”
4. Context Overflow / Compaction
Every agent must handle “context too long” and all use the same basic strategy:
- Detect: token count approaching context window (typically 80% threshold)
- Summarize: call LLM to compress conversation history into a summary
- Replace: swap history with summary + continuation marker
- Resume: next turn sees summary instead of full history
5. Streaming + Events
Every agent streams LLM output token-by-token and emits typed events:
turn_start/turn_endmessage_start/ text chunks /message_endtool_call/tool_resultusage(token counts)error
6. Abort/Cancel Signal
Every agent has a cooperative cancellation mechanism:
- Pi:
AbortSignalchecked between steps - Goose:
CancellationTokenchecked in loop - Kimi Code:
signal.throwIfAborted()at loop boundary - Grok Build: token-based cancellation
7. Session Persistence
Every agent persists conversation state across turns:
- Codex: thread/session JSONL files
- Goose:
SessionManagerwith SQLite - Kimi Code:
wire.jsonlrecords - Grok Build: SQLite journal
- Pi: session JSONL entry files
- OpenCode/Kilocode: Effect-TS database layer (Drizzle + SQLite)
Universal System Prompt Patterns
Despite different wording, every agent’s system prompt contains these same sections:
| Section | What it says | Present in |
|---|---|---|
| Identity | “You are X, a coding assistant” | All |
| Autonomy | “Keep going until done” | All |
| Tool preference | “Use specialized tools over bash” | All except Goose (delegated) |
| File editing | How to edit files (patch/edit/replace) | All |
| Output style | Be concise, use markdown | All |
| Project instructions | Read and obey AGENTS.md / project files | All |
| Verification | Run tests after changes | All |
| Safety | Don’t break things, confirm risky actions | All |
Universal Tool Categories
Despite different names, every agent provides tools in these categories:
| Category | Purpose | Examples |
|---|---|---|
| File Read | Read file contents | read, read_file, cat |
| File Write/Edit | Modify files | edit, write, apply_patch, search_replace |
| Shell | Execute commands | bash, shell, run_commands |
| Search | Find in codebase | grep, rg, glob, search_codebase |
| Directory | List files | ls, list_dir, glob |
These 5 categories are present in every agent. Everything else (web, plan, memory, subagent, workflow) is bonus.
File System as Architecture
File reading and writing aren’t just “tools” — they define the agent’s fundamental relationship with code. The read/write strategy determines:
- What the LLM sees (line numbers? anchors? diffs? raw content?)
- How edits are specified (exact match? line numbers? anchors? whole-file?)
- What can go wrong (stale references, ambiguous matches, merge conflicts)
The Read→Edit Coupling (present in ALL agents)
Every agent couples its read format to its edit format. The read tool produces output that the edit tool consumes as addressing:
| Agent | Read Format | Edit Addressing | Edit Mechanism |
|---|---|---|---|
| Codex | raw (via shell) | Context lines (3 before/after) | apply_patch — custom diff (@@-anchored hunks, +/- lines) |
| Cline | LINE_NUMBER→CONTENT | exact old_text match OR insert_line | edit_file (search/replace) + apply_patch (diff grammar) |
| Goose | (via MCP extensions) | (extension-dependent) | (extension-dependent) |
| Grok Build | LINE:HASH:CTX→CONTENT | Anchor-based (22:abc:rst) | Hashline edit (atomic batch, stale = reject all) |
| Grok Build (alt) | LINE_NUMBER→CONTENT | exact old_string match | search_replace (find & replace) |
| Kimi Code | (dynamic) | (dynamic) | (dynamic) |
| OpenCode/Kilocode | raw with line numbers | exact oldString match | edit (search/replace with replaceAll option) |
| Pi | raw with line numbers | exact old_text match | edit tool |
| Qwen Code | cat -n format (line numbers) | exact old_string match | edit (search/replace) |
Three Families of Edit Strategy
1. Exact String Match (most common)
- Used by: Cline, OpenCode, Kilocode, Qwen Code, Pi, Grok Build (search_replace)
old_stringmust match exactly once in the file → replaced withnew_stringold_string = ""ornull→ create new file- Failure mode: ambiguous match (appears 0 or >1 times)
- Mitigation:
replaceAllflag, or “add surrounding lines to make unique”
2. Diff/Patch Language
- Used by: Codex, Cline (apply_patch)
- Custom mini-language with
*** Begin Patch,@@ context,+/-lines - Addresses via context (like git diff) not exact match
- Failure mode: context doesn’t match current file state
- Advantage: can express multi-hunk changes in a single tool call
3. Anchor-Based (unique to Grok Build)
- Each line gets a content-derived hash anchor:
LINE:HASH→CONTENT - Edits reference anchors, not line numbers or string matches
- Atomic batch semantics: if any anchor is stale, ALL edits rejected
- Stale anchor → error response includes fresh anchors → model retries immediately
- Three scheme candidates: ContentOnly, ChunkFingerprint, CheckpointChain
- Advantage: robust to concurrent edits / line shifts
- Disadvantage: anchor churn after edits, more complex protocol
The File Creation Pattern
Every agent needs to handle “create new file” distinctly from “edit existing file”:
- Codex:
*** Add File: <path>header in patch language - Cline/OpenCode/Qwen:
old_string = null/empty+new_string = content - Grok Build (search_replace):
old_string = ""creates file - Pi: separate
writetool for full-file writes
Why This Is Architectural
The read/edit coupling shapes the entire agent experience:
- Token efficiency: Codex sends minimal context (3 lines); hashline requires anchors for every line read
- Reliability: exact-match can fail on repeated code; anchors handle line shifts gracefully
- Multi-edit atomicity: patch language batches hunks; search_replace is one-at-a-time; hashline batches are atomic
- Model burden: exact-match requires the model to reproduce code perfectly; patch format allows context-based targeting
- Read-before-write requirement: Nearly all agents enforce “must read before edit” (OpenCode errors if not, Grok Build’s prompt says “Read the file first”)
The Filesystem as Infrastructure
The filesystem isn’t just a target of code edits — it’s a core piece of the agent’s own infrastructure. Every agent uses files for three internal purposes beyond code editing:
A. Session Persistence (conversation state across turns)
Every agent persists conversation history to disk so sessions survive restarts:
| Agent | Storage Format | Location |
|---|---|---|
| Codex | JSONL (append-only) | session-{id}.jsonl |
| Goose | SQLite (v15 schema) | ~/.config/goose/sessions.db |
| Grok Build | SQLite (WAL/journal) | Session directory |
| Kimi Code | JSONL (wire.jsonl) | Per-agent file, append + rewrite |
| Pi | JSONL (append-only tree) | Per-session .jsonl file |
| OpenCode/Kilocode | SQLite (via Drizzle + Effect) | Database layer |
| Qwen Code | JSONL + output files | Session directory |
The two strategies:
- JSONL (Codex, Kimi Code, Pi, Qwen Code): Append-only log of records. Simple, streamable, easy to replay. Pi adds tree structure (parentId) for branching.
- SQLite (Goose, Grok Build, OpenCode): Structured queries, schema migrations, concurrent access. Goose is at schema v15 — it evolves.
B. Large Output Spilling (keeping context manageable)
When a tool produces output too large for the context window, every agent writes it to a temp file and replaces it with a pointer:
| Agent | Threshold | Strategy | File Location |
|---|---|---|---|
| Codex | ~2,500 tokens | Spill to file, show head/tail preview + path | <temp>/hook_outputs/<thread_id>/ |
| Goose | 200,000 chars | Write full output to file, replace with file path message | goose_mcp_response_*.txt (tempfile) |
| Qwen Code | 30,000 chars (shell) | head 1/5 + tail 4/5, save full to .output file | <projectTemp>/<toolName>.output |
| Grok Build | (configurable) | Per-tool truncation with output files | Session temp directory |
The universal pattern:
if output.size > threshold:
file = write_to_temp(full_output)
model_sees = f"Output too large ({size}). Saved to: {file}\n{head}...[TRUNCATED]...{tail}"
This creates a feedback loop with file reading: the model can then use the read tool to examine the spilled file if it needs the full content. The filesystem becomes a working memory extension.
C. Compaction Persistence (surviving context resets)
Some agents persist compaction artifacts beyond the current context:
- Grok Build: References
/tmp/compaction/segment_*.mdand/tmp/compaction/INDEX.mdas “out-of-band memory channels for a future work agent.” - Kimi Code: Plan versions persisted as
agents/<agentId>/plan/<planId>/v<N>.mdwith SHA256 content hash. Cold rebuild fromwire.jsonl. - Codex: Session logs enable resume from any point.
- Pi: Compaction entries stored in the session JSONL with file operation tracking (which files were read/modified).
D. Background Task Output
When agents run long-running commands in the background, the filesystem bridges the gap:
- Qwen Code (
shell.ts:3058):shell-${shellId}.output— background shell stdout streamed to file, model reads later viatask_outputtool - Codex: Hook outputs spilled to disk per-thread
- Goose: Scheduled recipe execution results in sessions
Why This Matters Architecturally
The filesystem serves as the agent’s external memory hierarchy:
┌─────────────────────────────────────────────────┐
│ L1: Context Window (fast, limited, volatile) │
├─────────────────────────────────────────────────┤
│ L2: Spilled Output Files (temp, read-on-demand) │
├─────────────────────────────────────────────────┤
│ L3: Session Persistence (durable, replayable) │
├─────────────────────────────────────────────────┤
│ L4: Project Files (the actual codebase) │
└─────────────────────────────────────────────────┘
- L1 → L2: Overflow. When tool output exceeds context budget, spill to file. Model can read back if needed.
- L1 → L3: Checkpoint. On compaction or session save, persist full state so it can be restored.
- L2 → L1: Recovery. Model uses read tool to pull spilled content back into context when needed.
- L4 → L1: The normal read-file-into-context flow for code editing.
This is functionally the same architecture as CPU cache hierarchies — the context window IS the L1 cache, and the filesystem is everything slower but larger.
The Conversation Shape
Every agent uses the same message shape for LLM communication:
[
{ role: "system", content: [assembled system prompt] },
{ role: "user", content: "user's request" },
{ role: "assistant", content: "thinking + tool_calls" },
{ role: "tool", content: "tool results" },
{ role: "assistant", content: "more tool_calls or final response" },
...
]
The only variation is whether thinking/reasoning is a separate content block or inline.
Key Structural Insight
The entire architecture can be reduced to:
An agentic coding CLI is a loop that repeatedly calls an LLM with (system_prompt + history + tool_definitions), executes any tool_calls in the response, appends results to history, and repeats until the LLM produces a text-only response or a guard fires.
Everything else — compaction, subagents, plan mode, skills, permissions, memory — is optimization or UX layered on top of this core loop.