Chapter 1What you just joined
DeepSeek Harness — dsh on the command line — is an agent harness. That is
the software that sits between a large language model and the real world: it builds the prompt, calls the
model, streams back the answer, runs the tools the model asks for (reading files, running shell commands,
searching the web), records everything that happened, and renders it for a human. The model supplies the
thinking. The harness supplies everything else — and "everything else" is a surprisingly large amount of
software.
Most harnesses are one program with extension points bolted on. This one inverts that. There is no core program. There is a small plugin framework, and then every single feature is a plugin — the model adapter, the tool registry, the conversation log, the permission system, the web interface, and even the main agent loop itself. What ships as "the product" is a list of 78 configuration rows saying which plugins to load.
That is the one idea you need to hold onto. Everything in this handbook is a consequence of it.
Because agent harnesses change constantly. New model APIs, new sandboxing rules, a different UI, a customer who needs execution to happen on a remote machine instead of a laptop. If those are code paths inside one program, every change is surgery on shared code. If they are plugins behind stable interfaces, a change is: write a new plugin, swap one line of configuration.
The price is that you cannot understand this codebase by reading it top to bottom. There is no top. You understand it by learning the seams — and that is what Part II teaches.
Who builds it
A small team inside DeepSeek — 28 people appear in the commit history available to this report, with roughly 8–10 committing regularly. Chapter 11 breaks down who works where.
One thing to know up front, from CONTRIBUTING.md: the project does not accept external pull requests right now. It is open source, and the team actively wants an ecosystem, but they want that ecosystem to live in other people's plugin repositories rather than in this tree. The document says so directly:
DeepSeek Harness is designed to be deeply customizable. We do not believe that packages in the official repository are inherently more important than packages created by the community. You may consider this repository an idea, an official showcase, and a source of inspiration, but not a mandate from us.
CONTRIBUTING.md
This matters for you practically: the plugin interfaces are treated as public API even though the version number starts with a zero. When you design something, assume a stranger will build on it.
What state the project is in
Version 0.1.0-rc.5, published to npm, pre-1.0. The root
AGENTS.md opens with a
section that will not survive the first real release:
Pre-release stance: foundation over blast radius. With no external consumers, prefer the correct foundation over compatibility shims: rename or repackage freely and update every reference together.
AGENTS.md
Take that literally. In the few weeks of history this report can see, whole package groups were renamed
— bash/ became shell/, self-modification/ became
extensions/, support/ became test-support/. If you find a document
or a memory that references an old path, the path moved and nobody left a redirect. That is deliberate.
Chapter 2Day one: make it run
Before reading any more architecture, get the thing running. The mental model lands much faster when you have watched it work.
Prerequisites
- Node
^22.19.0 || >=24.0.0— older majors are not supported. - pnpm — this is a pnpm workspace; npm and yarn will not resolve it correctly.
- A
DEEPSEEK_API_KEYin a root.envfile, for anything that talks to a real model. Everything else runs without one.
# install the whole workspace
pnpm install
# run one task through the one-shot runner (needs a key)
pnpm dsh --profile headless "list the files in this directory"
# see the plugin tree your machine actually boots — no key needed
pnpm dsh --profile web --dump-config
# the demos
pnpm run demo:acp # automation server over the Agent Client Protocol
pnpm run demo:cordis # the agent modifies its own running plugin tree
Then the checks. Do not run the whole suite — see Chapter 13 for why, and for how to choose. For now, just confirm the basics work:
pnpm run typecheck # tsc across every package
pnpm run lint # oxlint
pnpm run test # vitest unit tests
pnpm run build # tsc emits lib/types, tsdown bundles runtime
- docs/architecture.md — 129 lines, the whole map. Read it twice.
- docs/cordis-primer.md — the plugin framework in five ideas.
- docs/cordis-tutorial/ — seven hands-on chapters. Actually type them out; it takes an afternoon and it is the fastest path in.
- docs/glossary.md — this team uses precise words. Turn, step, and round are three different things and people will assume you know which is which.
- AGENTS.md —
the conventions. Skim now, return often.
CLAUDE.mdis a symlink to it.
Chapter 3The map of the repo
Nine top-level directories. Here is what each one is for.
Inside packages/, the 49 groups are not arbitrary. Each one is a capability family
— the definition of a capability plus its implementations plus the model-facing tool that uses it. Learn
these ten first; the rest follow the same pattern.
| Group | What lives there | You will care when… |
|---|---|---|
core/ | The spine: session log, prompt assembly, tool registry, agent types, and the loop. | Almost always. Start here. |
llm/ | Model vocabulary + DeepSeek adapters + retry + token metering. | Adding a provider or debugging a request. |
session/ | Durability: JSONL/SQLite persistence, projections, titles, telemetry. | Anything about saving or replaying conversations. |
fs/, shell/, subprocess/ | The file and command capabilities behind read, write, edit, bash. | Touching how the agent affects a machine. |
sandbox/ | Process confinement: bubblewrap, Landlock, Seatbelt. | Security work. |
interaction/ | Humans: approval prompts, permissions, slash commands, ask-user. | Anything a person clicks or confirms. |
subagent/ | Delegation to child agents — including other vendors' agents. | Multi-agent work. |
client/ | 39 packages: the entire browser UI, as plugins. | Front-end work. The busiest area in the repo. |
host/ | The server half of the web app: HTTP routes and the API gateway. | Front-end work that crosses the wire. |
bundle/ | The three shipped compositions: base, web-app, headless. | Changing what loads by default. |
Every group has a README.md listing its packages and their service keys, and
packages/README.md
indexes all of them. That file is the fastest way to find where something lives.
Chapter 4Everything is a plugin
The framework underneath is Cordis. It lives in vendor/ as pinned source
rather than as an npm dependency, because the team modifies it and refuses to have a version-skew problem
in the thing that boots the product.
Cordis has five ideas. Learn them and 80% of the repo becomes readable.
1. A plugin is an object
Either a plain object with an apply(ctx) function, or a class extending
Service. That is the whole definition.
2. A context is a repository of services
A service claims a stable key — ctx.tools, ctx.llm,
ctx.sessions — and everyone else finds it by key, never by importing the
implementation. This is why a provider can be swapped without touching its consumers.
3. Dependencies are declared, not ordered
A plugin lists what it needs in inject and Cordis waits until those services exist. Nobody
maintains a boot sequence.
4. Events are typed, and their dispatch mode is part of the contract
Four modes: emit (fire and forget), waterfall (middleware that can transform
or short-circuit), parallel, and serial. Plugins add their own event types by
TypeScript declaration merging, without editing the package that owns the event.
5. Registrations are reversible effects
Everything you contribute — a tool, a prompt section, an adapter, a listener — is installed through
ctx.effect() or ctx.on(), which return a disposer. Unload the
plugin and every contribution unwinds automatically.
What a plugin actually looks like
Here is a complete, working plugin — a tool the model can call. This is taken from the tool cookbook and it is genuinely all the code required:
import { readFile } from 'node:fs/promises'
import type { Context } from '@deepseek-ai/cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'
export const name = 'my-tool'
export const inject = ['tools'] // wait for ctx.tools to exist
export function apply(ctx: Context) {
ctx.tools.register(defineTool({
name: 'read_file',
description: 'Read a file from disk.', // what the MODEL sees
parameters: {
path: { type: 'string', required: true, description: 'Absolute path' },
limit: { type: 'number' }, // optional by default
},
output: {
schema: { type: 'string' },
render: (_args, value) => [{ type: 'text', text: value }],
},
async execute(args, exec) {
// args is TYPED from the schema above: { path: string; limit?: number }
// and already validated — you do not parse model JSON by hand
return readFile(args.path, { encoding: 'utf8', signal: exec.signal })
},
}))
}
Read that again with the five ideas in mind. inject declares the dependency.
ctx.tools is found by key. register() returns a disposer, so disposing this
plugin's fiber unregisters the tool. And the schema flows into system-prompt assembly automatically —
nobody wires it up.
Function plugins export name / inject / Config /
apply as named exports and must have no default export. Service packages do
the opposite: they default-export their service class.
Mix the two forms and the Loader silently discards the plugin's namespace — including its
inject — and your plugin loads before its dependencies exist, failing in a confusing way far
from the cause. This has its own numbered postmortem:
postmortem 0001.
Read it now so you recognise the symptom later.
Chapter 5How the app is assembled
There is no main() that wires the product together. Instead there is a YAML file listing
plugins. Here is a real fragment from the one-shot runner in
examples/headless-agent:
# User-settings document ($DSH_HOME/settings.yaml, hot-reloaded)
- id: settings
name: '@deepseek-ai/dsh-settings-file'
# Credential store: live process env over $DSH_HOME/.credentials.yaml
- id: credentials
name: '@deepseek-ai/dsh-credentials-local'
# The DeepSeek adapter. Swap to dsh-llm-pi-ai for the pi-ai-backed twin.
- id: llm-deepseek
name: '@deepseek-ai/dsh-llm-deepseek'
config:
thinking: enabled
reasoningEffort: max
models:
- id: deepseek-v4-pro
contextWindow: 128000
- id: subprocess
name: '@deepseek-ai/dsh-subprocess-local'
- id: bash
name: '@deepseek-ai/dsh-bash-local'
Each row has an id (a stable handle), a name
(the npm package), and optional config. Row order carries no meaning —
activation is driven by service availability, i.e. by inject.
Profiles, bundles, and patches
Real deployments do not hand-write the full list. They stack layers:
id and replaces its whole config, or inserts new rows.Patching a row swaps its entire config object. There is no deep merge. The base
bundle's own header comment explains the consequence: a row whose value differs between web and headless
deliberately does not live in base. Each mode bundle restates its complete
configuration, so any single row is only ever touched by one bundle layer plus the user's.
When you add configuration, ask yourself which layer owns it. Getting this wrong produces a setting that mysteriously reverts.
The command you will use constantly:
pnpm dsh --profile web --dump-config
When something is not loading, this is the first thing to run. It shows you the resolved tree, so you can see whether your plugin is even in it before you start debugging why it is not working.
Chapter 6The loop, step by step
Now the part everyone wants to understand: what actually happens when a user sends a message. The driver lives in packages/core/agent-loop/src/agent.ts and it is about 1,600 lines including its siblings — small, because all the policy lives in plugins.
The vocabulary (get this right)
| Word | Means |
|---|---|
| Step | One model request, plus the tool calls its response triggered. |
| Turn | One drain of admitted input. Contains zero or more steps. Opens before the first input is claimed, closes when nothing is owed. |
| Round | An outer policy iteration that contains a turn — a goal round, one Ralph attempt. Round counters belong to that policy, not to the session. |
A turn with zero steps is not a bug. If a hook rejects the input, the turn still opens and closes so the log records that something was attempted and blocked.
The inbox: two lanes and a wakeup bit
Input does not go straight to the model. It goes into the agent's inbox, which holds two ordered lists. There is one primitive:
send(message: UserMessage, target: 'next-turn' | 'next-step', wakeup: boolean)
and three named presets over it that you will see everywhere:
| Method | Lands in | Wakes the agent? | Used for |
|---|---|---|---|
followup() | next-turn | Yes | A normal user message. Gets its own turn. |
steer() | next-step | Yes | "Actually, stop and do this instead." Consumed at the next step boundary. |
inject() | next-step | No | Background context. Rides along with the next request but never starts one. |
That third row is the one people miss. inject() on an idle agent does nothing visible —
the context sits in the inbox until something else wakes the driver. That is intentional, and it is the
right behaviour for things like "the file you were editing changed on disk."
The inbox is also durable: every mutation is recorded as an
agent/inbox/spliced event, so pending work survives a reload and the UI rebuilds the queue from
the log rather than from memory.
One turn, end to end
Here is the actual code for the inner part of a step, lightly trimmed. Notice how little it does:
const stream = preparedCall?.stream(request) ?? this.loopCtx.llm.stream(request)
for await (const chunk of stream) {
signal.throwIfAborted()
// every chunk is appended to the durable log AS IT ARRIVES
chunkSeqs.push(this.session.append('assistant/chunk', { turn, step, chunk }).seq)
assembler.push(chunk)
}
const finish = assembler.finish
if (finish.kind === 'error' || finish.kind === 'aborted') {
// the loop does NOT know how to retry. it asks.
const action = await this.dispatch.waterfall('agent/request-error', { … })
if (action?.kind !== 'retry') throw new LlmError(…)
continue // a listener said retry — go around again
}
Retry is not in the loop. It is dsh-llm-retry, a plugin listening on
agent/request-error. So is recovery from an over-long context:
compaction-basic listens on the same event, compacts the history, and returns
retry.
This is the pattern you should copy. When you are asked to add behaviour, your first question is "which existing event should hear about this?" — not "where in the loop do I add an if?" Changing the loop requires updating docs/architecture.md, and that is a deliberate speed bump.
Chapter 7The log you must not break
If you remember one technical thing from this handbook, make it this one.
A session is an append-only log of typed events. The conversation history sent to the model is not stored anywhere — it is derived from the log every time, by folding a pure function over it.
The projection rule is one pure function, and its JSDoc explains why it must stay pure:
/**
* Project a single event into the LLM message it derives to, or null when it
* produces none … This is THE per-node projection rule: `Session.deriveMessages`
* folds it over the live surface, external reconstructors and pure projections
* fold the same function over a log prefix's surface to rebuild the exact
* messages any request was built from.
*/
export function deriveEventMessage(event: SessionEvent): Message | null {
switch (event.type) {
case 'user/message': return event.data
case 'assistant/message': return event.data.message.content.length === 0
? null // empty = usage-only, skip
: event.data.message
case 'tool/result': return event.data.message
default: return null // boundaries, chunks: trace data
}
}
Anything that reaches a model request must be reconstructable from the session log. A
runtime invariant asserts this. In practice it means: if you want to add a new kind of thing the model can
see, you cannot just append a string to the prompt. You must add a new event type to
SessionEventMap (by declaration merging, from your own package — about twenty packages do
this) and render it from the log.
The reason is debugging. If the model saw something the log does not contain, then nobody — not you, not a replay, not a bug report — can ever reconstruct why it behaved the way it did.
One more strictness you should expect: reading is fail-closed. A build that encounters
an event type it does not know refuses the log, unless the event was explicitly marked
ignorable: true by its producer. That feels harsh until you realise the alternative is silently
showing a user a different conversation than the model actually saw.
Chapter 8Tools
Tools are how the model does anything at all. There are 52 of them shipped across 24 packages —
bash, read, write, edit, grep,
glob, web_search, todo_write, subagent,
terminal_*, session_*, and more. The full generated list is in
docs/tool-catalog.md.
You saw the minimal shape in Chapter 4. What you did not see is everything the
registry does around your execute() function.
Three things about this that are easy to get wrong
Guards are not the same as the pre-execute waterfall. A waterfall runs listeners in registration order, and registration order depends on composition — which a user can change with a patch. That is fine for annotation and unacceptable for security. So there is a second layer: guards, which can only deny or abstain, never allow. Because they cannot allow, their order cannot change the outcome. If you are writing policy that must hold, write a guard.
Only three fields ever reach the model. When the registry builds tool schemas for a
request it uses an explicit allowlist: name, description,
parameters. Your timeoutMs, isConcurrencySafe, output schema, and UI
presenters are internal and must never leak into a prompt.
Presenters must be pure. presentCall(args) and
presentResult(args, result) decide how your tool renders as a card in the UI. They run during
live streaming and during log replay, so they may do no I/O, read no session state, and use no clock
or randomness. If you find yourself wanting the file's previous contents inside presentCall,
stop — that belongs in durable result metadata.
There is a reserved tool called run_code. Instead of calling tools one at a time, the model
writes the body of an async TypeScript function and calls await tools.read_file({ path })
directly. Argument and return types are generated from the same schemas you already wrote.
You do not integrate with this. Any tool you register is automatically available. The important property: those nested calls re-enter the complete pipeline above — permission denials come back as catchable rejections inside the program, not as text the model has to parse.
Chapter 9Capability seams
When the team says "add a capability," they mean a specific three-part structure. Getting this vocabulary right will make code review go much better.
packages/shell/ is the canonical example. Consumers depend on the
definition, never on a concrete provider, which is what keeps providers swappable.There are 55 service keys in the system. Here are the ones you will meet first:
| Key | Capability | Notable providers |
|---|---|---|
ctx.llm | Model adapters | llm-deepseek, llm-pi-ai, llm-replay (tests) |
ctx.sessions | The session store and log | core — one implementation |
ctx.tools | Tool registry + guarded pipeline | core — one implementation |
ctx.shell | Command execution | bash-local, bash-sandbox, pwsh-* |
ctx.fs | Filesystem access | fs-local, fs-sandbox, fs-e2b |
ctx.subprocess | Process spawning | subprocess-local, subprocess-e2b |
ctx.sandbox | Confinement | sandbox-local (bwrap / Landlock / Seatbelt) |
ctx.subagents | Delegation to child agents | 8 providers — see below |
ctx.sessionPersistence | Durability | JSONL, SQLite |
ctx.approval | Human confirmation | per front end |
Filesystem and subprocess providers share one execution world. So pointing those two at a remote
sandbox moves bash, the PTY terminals, and the language-server integration with them — no
forks of any tool. The E2B remote-sandbox proof of concept is just three packages:
e2b, fs-e2b, subprocess-e2b.
This is the test of whether a seam is real. When you design one, ask: "if someone swapped the provider, would everything downstream keep working?"
Chapter 10The rest of the system
You do not need these on day one, but you should know they exist and roughly where they live, so you recognise the names in review and in the event matrix.
| Subsystem | What it does | Where |
|---|---|---|
| Scope | Per-agent worlds. A tool or prompt section can be global or owned by exactly one agent; a scoped one shadows its global twin. tools.restrict() filters the global set per agent — and a filtered-away tool is absent from the prompt and refuses to execute. | core/scope |
| Presets | Compose a whole agent from a preset cordis.yml, per session. | preset/ |
| Subagents | Delegation behind one interface. Providers include a fresh in-process child, a fork seeded from the parent's history, an out-of-process child over ACP or the SDK, a real Codex app-server, and a real Claude Code child via the official Agent SDK. | subagent/ |
| Hooks | Bridges that read an existing Claude Code or Codex hooks.json and run those shell hooks faithfully, mapped onto this harness's own interception points. A "native hook" here is just an ordinary plugin. | hooks/ |
| Compaction | Shrinks history when the context fills — as a replace on the log surface, never a rewrite. | compaction/ |
| Spill | Oversized tool output goes to a file; the model gets a bounded preview plus a retrieval locator. | spill/ |
| Jobs | Background work with job_list / job_output / job_kill control tools. | jobs/ |
| Session query | Search and trace across past sessions, including SQLite full-text search, exposed to the model as five session_* tools. | session-query/ |
| Skills | Loadable instruction packs the model can pull in on demand. | skill/ |
| Plan / Goal / Schedule | Plan mode as logged state; durable same-session objectives; session-local scheduled follow-ups. | plan/, goal/, schedule/ |
| Workflow / Ralph | A workflow engine on worker threads, plus the "Ralph loop" — repeated fresh-agent attempts at one fixed objective. | workflow/ |
| Web GUI | 39 client packages registering into UI slots, plus the host half serving them. Typed RPC between the halves is generated by Typert. | client/, host/, typert/ |
| Self-modification | cordis_define / cordis_run / cordis_undefine: the model writes and mounts plugins into the runtime it is running inside, evaluated in a node:vm sandbox. | extensions/ |
| SDKs & protocols | JSON-RPC protocol + TypeScript client, an ACP automation server, MCP support, and a Python SDK that drives a bundled runtime binary. | sdk/, acp/, mcp/, python/ |
| Invariants | Every package registers runtime checks under its own npm name. See Chapter 14. | runtime-diagnostics/ |
Chapter 11Who works on what
There is no CODEOWNERS file in this repository. What follows is derived from git history, so treat it as "who has been active where recently," not as a formal ownership map. Confirm with your lead before assuming someone is a reviewer.
Two caveats on the method. First, this clone is shallow: it covers roughly
2026-07-20 to 2026-08-14 — about four weeks and 863 non-merge commits. Longer-tenured
ownership will not show up. Second, raw git log per directory is badly misleading here,
because release commits touch 222 files across 52 areas at once and make one person look like they own
everything. The numbers below exclude repo-wide sweeps (any commit touching more than
six areas or sixty files), which removes 174 commits and leaves 689 focused ones.
By area — who to ask
| Area | Focused commits | Most active, in order |
|---|---|---|
Web UI (packages/client) | 176 | imccyu, _Kerman, Yichen Jiang, creatixchu, ZiyaZhang |
Web app (apps/web) | 110 | Yichen Jiang, _Kerman, imccyu, creatixchu |
Documentation (docs/) | 143 | Turtle, Yichen Jiang, Tianyi Cui, imccyu |
Gates & generators (scripts/) | 123 | imccyu, Turtle, Yichen Jiang, Tianyi Cui |
CLI (apps/cli) | 57 | Turtle, Yichen Jiang, Huanqi Cao, imccyu |
| Examples | 46 | Hypatia May, Yichen Jiang, pku-xht, imccyu |
Web host (packages/host) | 37 | imccyu, _Kerman, Tianyi Cui, ZiyaZhang |
Core spine (packages/core) | 30 | Chinesezjc, j-xiang, Yichen Jiang, imccyu |
Bundles (packages/bundle) | 30 | Turtle, Tianyi Cui, imccyu, Huanqi Cao |
| Subagents | 29 | Hypatia May, pku-xht, j-xiang, imccyu |
Boot glue (packages/boot) | 18 | Turtle, Huanqi Cao, Tianyi Cui |
CI (.github/) | 18 | Chinesezjc, imccyu, Yichen Jiang |
| Python SDK | 14 | Yichen Jiang, imccyu, j-xiang, _Kerman |
| Agent presets | 14 | Yichen Jiang (dominant) |
| LLM adapters | 13 | j-xiang, Yichen Jiang, imccyu |
| Sandbox | 12 | Tianyi Cui, Huanqi Cao |
API gateway (packages/api) | 12 | imccyu (dominant) |
Self-modification (extensions/) | 12 | imccyu (dominant) |
| Feedback | 11 | Chinesezjc, ZiyaZhang, Turtle |
| Vendored Cordis | 11 | imccyu, Turtle |
| Typert (RPC codegen) | 6 | imccyu (sole) |
| MCP | 6 | Tianyi Cui (dominant) |
| Native Landlock addon | 4 | imccyu (sole) |
By person — what each has been building
| Contributor | Focused commits | Their patch of the map |
|---|---|---|
| imccyu | 122 | The broadest range in the repo. Web UI and host, the gate scripts, the API gateway, Typert, the self-modification toolset, the vendored framework, the native addon — and the release process. |
| Yichen Jiang | 108 | Also very broad, weighted to product surface: the web app and UI, documentation, agent presets, the Python SDK. Writes more Agent Notes than anyone. |
| Turtle | 77 | Documentation, the CLI, the gate scripts, bundles and boot glue. If you have a question about how the app assembles itself or where a doc belongs, this is the trail to follow. |
| Tianyi Cui | 70 | Cross-cutting standards work: gates, docs, the agent skills, plus sandbox, MCP, and schedule. |
| Chinesezjc | 60 | The core spine — the largest single contributor to packages/core in this window — plus CI, feedback, and code-runtime. |
| _Kerman | 57 | Front end. Client packages and the web app, with some host-side work. |
| creatixchu | 35 | Front end: client, web app, host, and the attachment subsystem. |
| ZiyaZhang | 34 | Front end plus the request-context packages and docs. |
| Huanqi Cao | 33 | Sandbox, boot, the CLI, workflow, and gate scripts. Systems-leaning. |
| pku-xht | 26 | Subagents, examples, and client work. |
| Hypatia May | 19 | Examples and subagents — the runnable compositions the snapshot tests boot. |
| j-xiang | 6 | Small but deep: core, subagents, LLM adapters, and Chinese documentation. |
| Yif, NI0317, fz | 10 / 7 / 5 | Front end — client packages and the web app. |
| xjt | 7 | Documentation and the bilingual translation corpus. |
Two things the data tells you about how this team works
Everyone writes Agent Notes. .agents/notes/ is the single most-touched
directory in the repository — 249 focused commits from 18 different people, ahead of every package and
ahead of docs/. For nine of the top eleven contributors it is their #1 or #2 area. Design
records are not bureaucracy layered on top of the work here; they are a large part of the work.
Budget time for yours.
The front end is the busiest part of the product. packages/client plus
apps/web is 286 focused commits, more than twice the core spine, docs, or scripts. If you are
joining to work on "the agent," be aware that most day-to-day motion is in the browser half — and that the
browser half is built out of the same plugin machinery as everything else.
Before you open a pull request, run this on the files you changed. It beats any table:
git log --no-merges --format='%an' -- packages/<the-thing-you-touched> | sort | uniq -c | sort -rn | head
Chapter 12Your first change
A good first task is adding a tool, because it exercises nearly every convention in the repo without requiring you to understand the loop deeply. The guided version is docs/user/develop/basic/tool.md; the contract reference is docs/cookbook/adding-a-tool.md.
Here is the full checklist of what a non-trivial change ships. Missing items are the most common reason a review stalls.
| # | Deliverable | Why it is required |
|---|---|---|
| 1 | The package, named @deepseek-ai/dsh-<name>, in the right group, with a tsconfig.json referencing every workspace dependency. | Naming and layout are gated mechanically. |
| 2 | Unit tests under tests/ (never src/__tests__/), including an HMR-safety test — dispose the fiber, assert your contribution is gone. | Per-file 100% line coverage is the CI gate. Disposal is what makes the plugin model real. |
| 3 | A real-composition test — boot a test-only cordis.yml through the Loader and assert model-visible or durable output. | Hand-built ctx.plugin(...) suites do not catch the "green tests, broken product" failure. See postmortem 0001. |
| 4 | An ./invariant companion registered under your exact npm name. | Every package has one. If you have nothing to check, export an empty installer whose comment starts No runtime invariant: and explains why — a gate rejects unexplained empties. |
| 5 | A README with purpose, APIs, extension points, a Model Experience section, and ## Known Limitations and Deferred Work. | All three sections are separately gated. |
| 6 | An Agent Note in .agents/notes/, in the same PR. | Required for anything beyond a mechanical edit. This is the decision record. |
| 7 | A keyless snapshot scenario if the change is model-, protocol-, or human-visible. | Package tests explicitly do not substitute for the assembled transcript. |
| 8 | Docs updated in the same commit — affected READMEs and JSDoc. | "Docs accompany every code change" is a stated rule, not a nicety. |
An ./invariant companion is smaller than it sounds. Here is a real one, from
tool-todo, checking that every todo snapshot reaching the durable log is well-formed:
const PACKAGE_NAME = '@deepseek-ai/dsh-tool-todo'
export const name = 'tool-todo-invariant'
export const inject = ['invariants']
function validateTodos(value: unknown, fail: InvariantFailure): void {
if (!Array.isArray(value)) fail('todo/write todos must be an array')
const seen = new Set<string>()
for (const item of value) {
const { content, status } = item as Record<string, unknown>
if (seen.has(content)) fail(`todo/write repeats content …`)
if (!TODO_STATUSES.has(status)) fail(`todo/write carries unknown status …`)
}
}
export const apply = (ctx: Context) =>
Promise.resolve(ctx.invariants.register(PACKAGE_NAME, install))
Note what its JSDoc says it deliberately does not check — how many todos are
in_progress — because that is per-deployment config, and "a log written while parallel work was
allowed must still replay after a deployment tightens the policy." That is the level of care expected in
these files.
Branching and pull requests
- Develop on a branch; open a PR against
master. - Fill in the PR template.
- Labels: exactly one
kind/*, every materialarea/*, plus a native Issue Type. - Split independent changes into separate PRs. If a change builds on another, use GitHub's official stacked-PR feature rather than hand-managed bases.
- Rewrites use
--force-with-lease, never raw--force.
Chapter 13The gates
There are 124 scripts in scripts/ and a lot of them can fail your PR. The important thing
to internalise first is which ones to run locally.
This is stated policy, not a shortcut: "Never default to the full suite or repeat a passing check for commit or push. CI owns exhaustive coverage and the platform matrix."
Match the evidence to what you changed: focused tests for behaviour, snapshots for model or user
output, doc-sync for docs, build and hygiene for published paths, real-API e2e for provider
behaviour. There is a skill — dsh-pre-push-checks — whose whole job is choosing the smallest
sufficient set.
| Command | What it covers |
|---|---|
pnpm run test | Vitest unit tests. |
pnpm run test:coverage | The CI coverage gate — per-file 100% lines on packages/*/*/src. Not test. |
pnpm run test:snapshot | Keyless replay of real example compositions, diffing transcripts and re-persisted logs. Filter with -t <name>. |
pnpm run test:e2e | Real-API tests. Self-skips without DEEPSEEK_API_KEY. |
pnpm run test:web | Chromium browser snapshots. Required Linux PR gate. |
pnpm run typecheck / lint | tsc / oxlint. |
pnpm run doc-sync | All documentation gates at once — links, wrapping, budgets, translation pairing, catalog freshness, Mermaid, JSDoc. |
pnpm run hygiene | knip (dead exports) + publint + workspace constraints + NodeNext consumer check. |
pnpm run duplication | Cross-file clone detection. |
pnpm run build | tsc emits lib/types; tsdown bundles the runtime. |
Gates that surprise newcomers
- Per-file 100% coverage. Read an uncovered line as a question: is this code dead? The policy says outright that an uncovered line "is often dead code the gate is correctly flagging for deletion, not a missing test to bolt on."
- Documentation word budgets that ratchet down over time. Raising a ceiling requires justification in the PR.
- Bilingual pairing. Docs exist in English and Chinese with a paired
.i18n.yaml. A gate checks they stay in sync. - Type blocks in docs are verified against source. A
ts type-equivfence in a Markdown file is checked against the real type — so documentation cannot drift silently. - Generated catalogs must be fresh. The module graph, event matrix, capability-seam graph, config catalog, tool catalog, and persistence catalog are generated and CI verifies you regenerated them.
- One trailing newline per file, enforced by a pre-commit hook.
Chapter 14Rules that will bite you
These are the conventions from AGENTS.md and packages/AGENTS.md that are non-obvious. Each one exists because something went wrong once.
Waterfall listeners must call next()
A waterfall is around-middleware. Returning without calling next() short-circuits the whole
chain — every listener after you never runs. That is a legitimate design for a listener that owns a
decision, and a bug for one that only annotates. Know which you are writing.
Use ctx.get(name) for optional services
Reserve ctx.<name> for services you declared in inject. The property proxy
is topology-sensitive; ctx.get reads the global service store. Mixing them up produces
"service is undefined" in one composition and not another.
Enforce a decision where it is made
Leaving a tool out of the prompt is not enforcement. Neither is a facade, a wrapper, or listener order — anything a direct caller can bypass. Test denial through the executor. The corollary is that a filtered-away tool must be indistinguishable from a nonexistent one.
No hardcoded tunables in plugins
Anything that varies by deployment is a validated Config field settable from
cordis.yml. A DEFAULT_* constant or a test hook does not count as configurability.
Protocol constants, external specifications, and security invariants stay fixed.
Explicit beats implicit at package boundaries
Defaulting is an explicit resolve(request): Spec step in the owning implementation, never a
hidden ?? default buried in run(). The shell request/spec split is the template.
Misconfiguration fails loud
At load when self-contained, otherwise at the earliest resolvable point. Never silently skip a missing referent.
An empty catch must name what it swallows
…and why nothing else can reach it. Keep the try to a single statement.
Branded IDs, not bare strings
Opaque IDs that cross package boundaries use Branded<B>, so a SessionId
cannot be passed where a CallId is expected. Structurally they are strings; at the type level
they are not interchangeable.
Trust TypeScript at typed same-process boundaries
Do not add runtime validation or hostile-input tests for values the static interface already guarantees. Validate at the real boundaries: parsers, config, queues, model and tool JSON, files, workers, processes, and the wire. Defensive code everywhere is treated as noise here.
Write prose like a contract, not a transcript
This one catches almost everyone. Comments and docs state complete contracts and current behaviour —
never reasoning narration. No "this used to…", no "a later PR will…", no restating what the code plainly
does, no review history. Use direct, concrete terms; the standard explicitly asks you to write
response fields rather than response shape, and to reserve the word contract for
actual preconditions and guarantees. There is a skill, dsh-trim-cot-leakage, that exists purely
to audit for this.
A large share of this codebase is written with agent assistance, and reasoning-transcript prose is the characteristic failure mode of agent-written text. The rules — and the gate — are a deliberate counterweight. It also means the comments you read in this repo are unusually trustworthy: when a JSDoc block says something is invariant, it is describing a real obligation, not a passing thought.
Chapter 15Where to look things up
Much of this repo's reference material is generated from source and verified fresh in CI. When you need a fact, prefer these over grepping.
| Question | Generated answer |
|---|---|
| "What events exist, who fires them, who listens?" | docs/event-producer-consumer.md |
| "What services exist and who implements them?" | docs/capability-seams.md |
| "What config keys does this plugin accept?" | docs/config-catalog.md |
| "What tools ship, with what schemas?" | docs/tool-catalog.md |
| "What depends on what?" | docs/module-graph.md |
| "What gets written to disk?" | docs/persistence-catalog.md |
| "What does this word mean here?" | docs/glossary.md |
| "How does one turn actually flow?" | docs/agent-lifecycle.md · tool-execution-pipeline.md |
| "Why is it like this?" | .agents/notes/implemented/ — and .agents/notes/rejected/, which keeps 22 designs that were considered and turned down. |
That last row is worth a habit. The expensive question in a mature codebase is never "how does this work" — it is "why isn't it the obvious other way." The rejected notes answer that directly, and reading a few of them will teach you this team's taste faster than anything else.
- Everything is a plugin, found by service key, registered as a reversible effect.
- The app is assembled from config rows, in layers, patchable by
id. - The session log is the source of truth. Model history is derived, compaction shadows rather than rewrites, and model-visible means logged.
- New behaviour attaches to an event, not to a branch in the loop.
- A change is not done until it has tests, an invariant companion, a README, an Agent Note, and a snapshot if anything visible changed.