Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Agentic CLI Research

Exploring the architectures of 10 open-source coding agents to understand what makes them the same and what makes them different.

Agents under study: Codex (OpenAI), Cline, Kimi Code (Moonshot), Goose (Block), Grok Build (xAI), Kilocode, OpenCode, OpenHands, Pi, Qwen Code (Alibaba).

Research Files

Each file covers one architectural component. They are mutually exclusive — no content is repeated across files.

FileCovers
00-universal-architecture.mdThe shared skeleton: core loop, 7 universal components, conversation shape, filesystem as memory hierarchy
01-system-prompts.mdHow each agent defines identity, personality, and behavioral rules
02-tool-system.mdTool definition, registration, execution, parallelism, and dynamic loading
03-file-editing.mdThe read→edit coupling, three edit strategy families, file creation patterns
04-context-and-memory.mdCompaction strategies, session persistence, output spilling, token optimization
05-control-flow.mdPlan mode, permissions, subagents, orchestration, hooks
06-permission-safety.mdLLM classifiers, static rules, and safety classification across agents
07-workflow-orchestration.mdQwen Code workflows, Goose recipes, Grok Build goals, Kimi Code batch
08-multi-file-atomicity.mdRollback strategies, atomic edits, worktree isolation, error recovery
09-hashline-schemes.mdGrok Build’s three anchor schemes: ContentOnly, ChunkFingerprint, CheckpointChain
10-error-recovery.mdDoom-loop detection, circuit breakers, stall recovery
11-mcp-integration.mdMCP server discovery, lifecycle, transport, auth patterns
12-streaming-tui.mdRendering frameworks, streaming architecture, cancellation
13-agent-comparison.mdQuick-reference matrix and per-agent profiles

Key Finding

An agentic coding CLI is a loop that repeatedly calls an LLM with (system_prompt + history + tool_definitions), executes any tool_calls in the response, appends results to history, and repeats until the LLM produces a text-only response or a guard fires.

Everything else — compaction, subagents, plan mode, skills, permissions, memory — is optimization or UX layered on top of this core loop. The filesystem serves as the agent’s external memory hierarchy (L1=context window, L2=spilled output, L3=session persistence, L4=codebase).