Agentic CLI Research
Exploring the architectures of 10 open-source coding agents to understand what makes them the same and what makes them different.
Agents under study: Codex (OpenAI), Cline, Kimi Code (Moonshot), Goose (Block), Grok Build (xAI), Kilocode, OpenCode, OpenHands, Pi, Qwen Code (Alibaba).
Research Files
Each file covers one architectural component. They are mutually exclusive — no content is repeated across files.
| File | Covers |
|---|---|
| 00-universal-architecture.md | The shared skeleton: core loop, 7 universal components, conversation shape, filesystem as memory hierarchy |
| 01-system-prompts.md | How each agent defines identity, personality, and behavioral rules |
| 02-tool-system.md | Tool definition, registration, execution, parallelism, and dynamic loading |
| 03-file-editing.md | The read→edit coupling, three edit strategy families, file creation patterns |
| 04-context-and-memory.md | Compaction strategies, session persistence, output spilling, token optimization |
| 05-control-flow.md | Plan mode, permissions, subagents, orchestration, hooks |
| 06-permission-safety.md | LLM classifiers, static rules, and safety classification across agents |
| 07-workflow-orchestration.md | Qwen Code workflows, Goose recipes, Grok Build goals, Kimi Code batch |
| 08-multi-file-atomicity.md | Rollback strategies, atomic edits, worktree isolation, error recovery |
| 09-hashline-schemes.md | Grok Build’s three anchor schemes: ContentOnly, ChunkFingerprint, CheckpointChain |
| 10-error-recovery.md | Doom-loop detection, circuit breakers, stall recovery |
| 11-mcp-integration.md | MCP server discovery, lifecycle, transport, auth patterns |
| 12-streaming-tui.md | Rendering frameworks, streaming architecture, cancellation |
| 13-agent-comparison.md | Quick-reference matrix and per-agent profiles |
Key Finding
An agentic coding CLI is a loop that repeatedly calls an LLM with (system_prompt + history + tool_definitions), executes any tool_calls in the response, appends results to history, and repeats until the LLM produces a text-only response or a guard fires.
Everything else — compaction, subagents, plan mode, skills, permissions, memory — is optimization or UX layered on top of this core loop. The filesystem serves as the agent’s external memory hierarchy (L1=context window, L2=spilled output, L3=session persistence, L4=codebase).