Engineering Onboarding

The DeepSeek Harness Handbook

Everything a new engineer needs to be productive in deepseek-harness: what the product is, how the codebase is put together, what happens when a user sends a message, who works on which part, and how to land your first change without tripping over the rules.

219packages
198klines of src
231klines of tests
52model tools
60events
55services
28contributors
1,372agent notes
Part I · Arrival

Chapter 1What you just joined

DeepSeek Harness — dsh on the command line — is an agent harness. That is the software that sits between a large language model and the real world: it builds the prompt, calls the model, streams back the answer, runs the tools the model asks for (reading files, running shell commands, searching the web), records everything that happened, and renders it for a human. The model supplies the thinking. The harness supplies everything else — and "everything else" is a surprisingly large amount of software.

Most harnesses are one program with extension points bolted on. This one inverts that. There is no core program. There is a small plugin framework, and then every single feature is a plugin — the model adapter, the tool registry, the conversation log, the permission system, the web interface, and even the main agent loop itself. What ships as "the product" is a list of 78 configuration rows saying which plugins to load.

That is the one idea you need to hold onto. Everything in this handbook is a consequence of it.

Why they did it that way

Because agent harnesses change constantly. New model APIs, new sandboxing rules, a different UI, a customer who needs execution to happen on a remote machine instead of a laptop. If those are code paths inside one program, every change is surgery on shared code. If they are plugins behind stable interfaces, a change is: write a new plugin, swap one line of configuration.

The price is that you cannot understand this codebase by reading it top to bottom. There is no top. You understand it by learning the seams — and that is what Part II teaches.

Who builds it

A small team inside DeepSeek — 28 people appear in the commit history available to this report, with roughly 8–10 committing regularly. Chapter 11 breaks down who works where.

One thing to know up front, from CONTRIBUTING.md: the project does not accept external pull requests right now. It is open source, and the team actively wants an ecosystem, but they want that ecosystem to live in other people's plugin repositories rather than in this tree. The document says so directly:

DeepSeek Harness is designed to be deeply customizable. We do not believe that packages in the official repository are inherently more important than packages created by the community. You may consider this repository an idea, an official showcase, and a source of inspiration, but not a mandate from us.

CONTRIBUTING.md

This matters for you practically: the plugin interfaces are treated as public API even though the version number starts with a zero. When you design something, assume a stranger will build on it.

What state the project is in

Version 0.1.0-rc.5, published to npm, pre-1.0. The root AGENTS.md opens with a section that will not survive the first real release:

Pre-release stance: foundation over blast radius. With no external consumers, prefer the correct foundation over compatibility shims: rename or repackage freely and update every reference together.

AGENTS.md

Take that literally. In the few weeks of history this report can see, whole package groups were renamed — bash/ became shell/, self-modification/ became extensions/, support/ became test-support/. If you find a document or a memory that references an old path, the path moved and nobody left a redirect. That is deliberate.

Chapter 2Day one: make it run

Before reading any more architecture, get the thing running. The mental model lands much faster when you have watched it work.

Prerequisites

# install the whole workspace
pnpm install

# run one task through the one-shot runner (needs a key)
pnpm dsh --profile headless "list the files in this directory"

# see the plugin tree your machine actually boots — no key needed
pnpm dsh --profile web --dump-config

# the demos
pnpm run demo:acp      # automation server over the Agent Client Protocol
pnpm run demo:cordis   # the agent modifies its own running plugin tree

Then the checks. Do not run the whole suite — see Chapter 13 for why, and for how to choose. For now, just confirm the basics work:

pnpm run typecheck      # tsc across every package
pnpm run lint           # oxlint
pnpm run test           # vitest unit tests
pnpm run build          # tsc emits lib/types, tsdown bundles runtime
Your day-one reading list, in order
  1. docs/architecture.md — 129 lines, the whole map. Read it twice.
  2. docs/cordis-primer.md — the plugin framework in five ideas.
  3. docs/cordis-tutorial/ — seven hands-on chapters. Actually type them out; it takes an afternoon and it is the fastest path in.
  4. docs/glossary.md — this team uses precise words. Turn, step, and round are three different things and people will assume you know which is which.
  5. AGENTS.md — the conventions. Skim now, return often. CLAUDE.md is a symlink to it.

Chapter 3The map of the repo

Nine top-level directories. Here is what each one is for.

packages/ 219 npm packages, the entire product. Grouped by capability: packages/<group>/<package>/ → @deepseek-ai/dsh-<package> apps/ the two shipped applications: cli/ (terminal) and web/ (browser) examples/ runnable cordis.yml compositions — the real entry paths tests boot docs/ 215 Markdown files, bilingual (EN + 中文), several auto-generated .agents/ 1,372 Agent Notes (design decisions) + agent skills and workflows scripts/ 124 scripts: repo gates, verifiers, and doc/catalog generators vendor/ 9 pinned source copies of the Cordis framework — treat as read-only python/ Python SDK + a bundled runtime binary native/ the Landlock sandbox Node addon (Linux confinement) website/ VitePress projection of selected docs/ pages

Inside packages/, the 49 groups are not arbitrary. Each one is a capability family — the definition of a capability plus its implementations plus the model-facing tool that uses it. Learn these ten first; the rest follow the same pattern.

GroupWhat lives thereYou will care when…
core/The spine: session log, prompt assembly, tool registry, agent types, and the loop.Almost always. Start here.
llm/Model vocabulary + DeepSeek adapters + retry + token metering.Adding a provider or debugging a request.
session/Durability: JSONL/SQLite persistence, projections, titles, telemetry.Anything about saving or replaying conversations.
fs/, shell/, subprocess/The file and command capabilities behind read, write, edit, bash.Touching how the agent affects a machine.
sandbox/Process confinement: bubblewrap, Landlock, Seatbelt.Security work.
interaction/Humans: approval prompts, permissions, slash commands, ask-user.Anything a person clicks or confirms.
subagent/Delegation to child agents — including other vendors' agents.Multi-agent work.
client/39 packages: the entire browser UI, as plugins.Front-end work. The busiest area in the repo.
host/The server half of the web app: HTTP routes and the API gateway.Front-end work that crosses the wire.
bundle/The three shipped compositions: base, web-app, headless.Changing what loads by default.

Every group has a README.md listing its packages and their service keys, and packages/README.md indexes all of them. That file is the fastest way to find where something lives.

Part II · The machine

Chapter 4Everything is a plugin

The framework underneath is Cordis. It lives in vendor/ as pinned source rather than as an npm dependency, because the team modifies it and refuses to have a version-skew problem in the thing that boots the product.

Cordis has five ideas. Learn them and 80% of the repo becomes readable.

1. A plugin is an object

Either a plain object with an apply(ctx) function, or a class extending Service. That is the whole definition.

2. A context is a repository of services

A service claims a stable key — ctx.tools, ctx.llm, ctx.sessions — and everyone else finds it by key, never by importing the implementation. This is why a provider can be swapped without touching its consumers.

3. Dependencies are declared, not ordered

A plugin lists what it needs in inject and Cordis waits until those services exist. Nobody maintains a boot sequence.

4. Events are typed, and their dispatch mode is part of the contract

Four modes: emit (fire and forget), waterfall (middleware that can transform or short-circuit), parallel, and serial. Plugins add their own event types by TypeScript declaration merging, without editing the package that owns the event.

5. Registrations are reversible effects

Everything you contribute — a tool, a prompt section, an adapter, a listener — is installed through ctx.effect() or ctx.on(), which return a disposer. Unload the plugin and every contribution unwinds automatically.

The Context a repository of services, keyed by name ctx.tools · ctx.llm · ctx.sessions · ctx.shell … dsh-tool-bash registers a tool dsh-llm-deepseek registers an adapter dsh-hooks-codex registers a listener dsh-agent-loop consumes by key dsh-acp consumes by key ui-conversation consumes by key providers contribute consumers inject every contribution returns a disposer — unload the plugin and it unwinds cleanly
No plugin imports another plugin's implementation. They meet at a service key on the shared context, which is what makes any of them swappable from configuration.

What a plugin actually looks like

Here is a complete, working plugin — a tool the model can call. This is taken from the tool cookbook and it is genuinely all the code required:

a minimal tool plugin
import { readFile } from 'node:fs/promises'
import type { Context } from '@deepseek-ai/cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'

export const name = 'my-tool'
export const inject = ['tools']              // wait for ctx.tools to exist

export function apply(ctx: Context) {
  ctx.tools.register(defineTool({
    name: 'read_file',
    description: 'Read a file from disk.',      // what the MODEL sees
    parameters: {
      path:  { type: 'string', required: true, description: 'Absolute path' },
      limit: { type: 'number' },                  // optional by default
    },
    output: {
      schema: { type: 'string' },
      render: (_args, value) => [{ type: 'text', text: value }],
    },
    async execute(args, exec) {
      // args is TYPED from the schema above: { path: string; limit?: number }
      // and already validated — you do not parse model JSON by hand
      return readFile(args.path, { encoding: 'utf8', signal: exec.signal })
    },
  }))
}

Read that again with the five ideas in mind. inject declares the dependency. ctx.tools is found by key. register() returns a disposer, so disposing this plugin's fiber unregisters the tool. And the schema flows into system-prompt assembly automatically — nobody wires it up.

The one syntax trap

Function plugins export name / inject / Config / apply as named exports and must have no default export. Service packages do the opposite: they default-export their service class.

Mix the two forms and the Loader silently discards the plugin's namespace — including its inject — and your plugin loads before its dependencies exist, failing in a confusing way far from the cause. This has its own numbered postmortem: postmortem 0001. Read it now so you recognise the symptom later.

Chapter 5How the app is assembled

There is no main() that wires the product together. Instead there is a YAML file listing plugins. Here is a real fragment from the one-shot runner in examples/headless-agent:

examples/headless-agent/cordis.yml (abridged)
# User-settings document ($DSH_HOME/settings.yaml, hot-reloaded)
- id: settings
  name: '@deepseek-ai/dsh-settings-file'

# Credential store: live process env over $DSH_HOME/.credentials.yaml
- id: credentials
  name: '@deepseek-ai/dsh-credentials-local'

# The DeepSeek adapter. Swap to dsh-llm-pi-ai for the pi-ai-backed twin.
- id: llm-deepseek
  name: '@deepseek-ai/dsh-llm-deepseek'
  config:
    thinking: enabled
    reasoningEffort: max
    models:
      - id: deepseek-v4-pro
        contextWindow: 128000

- id: subprocess
  name: '@deepseek-ai/dsh-subprocess-local'

- id: bash
  name: '@deepseek-ai/dsh-bash-local'

Each row has an id (a stable handle), a name (the npm package), and optional config. Row order carries no meaning — activation is driven by service availability, i.e. by inject.

Profiles, bundles, and patches

Real deployments do not hand-write the full list. They stack layers:

applied in order, each layer may replace a row by id or insert new rows 1 · dsh-base 78 rows — model adapters, tools, persistence, sandbox, approval, settings, telemetry 2 · dsh-web-app (51 rows)  or  dsh-headless (3 rows) the mode bundle — this single choice is the difference between the two products 3 · the profile's own cordis.patch.yml 4 · the Harness-home patch 5 · any --patch overlay passed on the command line The plugin tree that actually boots dsh --dump-config prints exactly this every row it prints can be replaced by a patch of your own
A bundle is distributable config rows plus the code they mount. A profile is a named stack of bundles stored in the Harness home. A patch targets a row by id and replaces its whole config, or inserts new rows.
A patch replaces, it does not merge

Patching a row swaps its entire config object. There is no deep merge. The base bundle's own header comment explains the consequence: a row whose value differs between web and headless deliberately does not live in base. Each mode bundle restates its complete configuration, so any single row is only ever touched by one bundle layer plus the user's.

When you add configuration, ask yourself which layer owns it. Getting this wrong produces a setting that mysteriously reverts.

The command you will use constantly:

pnpm dsh --profile web --dump-config

When something is not loading, this is the first thing to run. It shows you the resolved tree, so you can see whether your plugin is even in it before you start debugging why it is not working.

Chapter 6The loop, step by step

Now the part everyone wants to understand: what actually happens when a user sends a message. The driver lives in packages/core/agent-loop/src/agent.ts and it is about 1,600 lines including its siblings — small, because all the policy lives in plugins.

The vocabulary (get this right)

WordMeans
StepOne model request, plus the tool calls its response triggered.
TurnOne drain of admitted input. Contains zero or more steps. Opens before the first input is claimed, closes when nothing is owed.
RoundAn outer policy iteration that contains a turn — a goal round, one Ralph attempt. Round counters belong to that policy, not to the session.

A turn with zero steps is not a bug. If a hook rejects the input, the turn still opens and closes so the log records that something was attempted and blocked.

The inbox: two lanes and a wakeup bit

Input does not go straight to the model. It goes into the agent's inbox, which holds two ordered lists. There is one primitive:

send(message: UserMessage, target: 'next-turn' | 'next-step', wakeup: boolean)

and three named presets over it that you will see everywhere:

MethodLands inWakes the agent?Used for
followup()next-turnYesA normal user message. Gets its own turn.
steer()next-stepYes"Actually, stop and do this instead." Consumed at the next step boundary.
inject()next-stepNoBackground context. Rides along with the next request but never starts one.

That third row is the one people miss. inject() on an idle agent does nothing visible — the context sits in the inbox until something else wakes the driver. That is intentional, and it is the right behaviour for things like "the file you were editing changed on disk."

The inbox is also durable: every mutation is recorded as an agent/inbox/spliced event, so pending work survives a reload and the UI rebuilds the queue from the log rather than from memory.

One turn, end to end

TURN  —  opens on turn/start, closes on turn/end  (both durable log events) claim next-step input + one queued next-turn message assemble prompt sections + tool schemas agent/pre-step  · waterfall 13 plugins listen · returns reject | enter(messages) reject → close the turn, spend no step the attempt is still recorded in the log STEP  — step/start … step/end append entered messages as user/message derive model history from the log llm/stream → assistant/chunk* → message tool/call* → the tool pipeline → tool/result* tools owe another request, or new next-step input arrived → loop again agent/turn-stopping · serial · last chance to continue
Nothing in this diagram is a branch inside the loop except the loop itself. Every named event is an extension point where a plugin attaches behaviour.

Here is the actual code for the inner part of a step, lightly trimmed. Notice how little it does:

packages/core/agent-loop/src/agent.ts — inside step()
const stream = preparedCall?.stream(request) ?? this.loopCtx.llm.stream(request)

for await (const chunk of stream) {
  signal.throwIfAborted()
  // every chunk is appended to the durable log AS IT ARRIVES
  chunkSeqs.push(this.session.append('assistant/chunk', { turn, step, chunk }).seq)
  assembler.push(chunk)
}

const finish = assembler.finish
if (finish.kind === 'error' || finish.kind === 'aborted') {
  // the loop does NOT know how to retry. it asks.
  const action = await this.dispatch.waterfall('agent/request-error', { … })
  if (action?.kind !== 'retry') throw new LlmError(…)
  continue   // a listener said retry — go around again
}
Why this matters to you

Retry is not in the loop. It is dsh-llm-retry, a plugin listening on agent/request-error. So is recovery from an over-long context: compaction-basic listens on the same event, compacts the history, and returns retry.

This is the pattern you should copy. When you are asked to add behaviour, your first question is "which existing event should hear about this?" — not "where in the loop do I add an if?" Changing the loop requires updating docs/architecture.md, and that is a deliberate speed bump.

Chapter 7The log you must not break

If you remember one technical thing from this handbook, make it this one.

A session is an append-only log of typed events. The conversation history sent to the model is not stored anywhere — it is derived from the log every time, by folding a pure function over it.

THE LOG — append-only, monotonic seq, nothing is ever deleted or edited turn/start user/message asst/chunk×N asst/message tool/call tool/result asst/message (replace) SURFACE — only message-producing events carry a surfaceOp marker append · append · append · replace → shadows the range behind it a compaction is an APPEND that hides earlier nodes — it never rewrites or deletes them What the MODEL sees deriveMessages() folds over the surface, so shadowed nodes are gone What the HUMAN sees append-origin events only so a compaction can never erase conversation the user already read two projections · one source of truth · both reproducible from a log prefix
This is the single most important structure in the codebase. Replay, forking, transcripts, telemetry, and the snapshot test suite are all consequences of it.

The projection rule is one pure function, and its JSDoc explains why it must stay pure:

packages/core/session/src/surface.ts
/**
 * Project a single event into the LLM message it derives to, or null when it
 * produces none … This is THE per-node projection rule: `Session.deriveMessages`
 * folds it over the live surface, external reconstructors and pure projections
 * fold the same function over a log prefix's surface to rebuild the exact
 * messages any request was built from.
 */
export function deriveEventMessage(event: SessionEvent): Message | null {
  switch (event.type) {
    case 'user/message':      return event.data
    case 'assistant/message': return event.data.message.content.length === 0
                                    ? null            // empty = usage-only, skip
                                    : event.data.message
    case 'tool/result':       return event.data.message
    default:                   return null   // boundaries, chunks: trace data
  }
}
The rule you will be held to: model-visible ⟺ logged

Anything that reaches a model request must be reconstructable from the session log. A runtime invariant asserts this. In practice it means: if you want to add a new kind of thing the model can see, you cannot just append a string to the prompt. You must add a new event type to SessionEventMap (by declaration merging, from your own package — about twenty packages do this) and render it from the log.

The reason is debugging. If the model saw something the log does not contain, then nobody — not you, not a replay, not a bug report — can ever reconstruct why it behaved the way it did.

One more strictness you should expect: reading is fail-closed. A build that encounters an event type it does not know refuses the log, unless the event was explicitly marked ignorable: true by its producer. That feels harsh until you realise the alternative is silently showing a user a different conversation than the model actually saw.

Chapter 8Tools

Tools are how the model does anything at all. There are 52 of them shipped across 24 packages — bash, read, write, edit, grep, glob, web_search, todo_write, subagent, terminal_*, session_*, and more. The full generated list is in docs/tool-catalog.md.

You saw the minimal shape in Chapter 4. What you did not see is everything the registry does around your execute() function.

model asks for a tool tool/call logged BEFORE execution tools/pre-execute · waterfall hook bridges · permission · sandbox ask → ctx.approval no answer = deny monotonic guards deny or abstain — never allow tools/execute around: timeout, retry, metrics your execute() body returns one canonical JSON value tools/post-execute → finalizeContent → tools/result spill oversized output · bound it · freeze it · log tool/result back to the loop, and to the UI as a rendered result card
Your tool body is the small box on the right. Everything else is policy contributed by other plugins — which is why tools stay simple and policy stays reusable across all 52 of them.

Three things about this that are easy to get wrong

Guards are not the same as the pre-execute waterfall. A waterfall runs listeners in registration order, and registration order depends on composition — which a user can change with a patch. That is fine for annotation and unacceptable for security. So there is a second layer: guards, which can only deny or abstain, never allow. Because they cannot allow, their order cannot change the outcome. If you are writing policy that must hold, write a guard.

Only three fields ever reach the model. When the registry builds tool schemas for a request it uses an explicit allowlist: name, description, parameters. Your timeoutMs, isConcurrencySafe, output schema, and UI presenters are internal and must never leak into a prompt.

Presenters must be pure. presentCall(args) and presentResult(args, result) decide how your tool renders as a card in the UI. They run during live streaming and during log replay, so they may do no I/O, read no session state, and use no clock or randomness. If you find yourself wanting the file's previous contents inside presentCall, stop — that belongs in durable result metadata.

Code Mode: your tool gets this for free

There is a reserved tool called run_code. Instead of calling tools one at a time, the model writes the body of an async TypeScript function and calls await tools.read_file({ path }) directly. Argument and return types are generated from the same schemas you already wrote.

You do not integrate with this. Any tool you register is automatically available. The important property: those nested calls re-enter the complete pipeline above — permission denials come back as catchable rejections inside the program, not as text the model has to parse.

Chapter 9Capability seams

When the team says "add a capability," they mean a specific three-part structure. Getting this vocabulary right will make code review go much better.

Service Definition dsh-shell → ctx.shell a Cordis Service, never a TS interface SERVICE PROVIDER dsh-bash-local SERVICE PROVIDER dsh-bash-sandbox SERVICE PROVIDER dsh-pwsh-local CONSUMER dsh-tool-bash The seam is all three roles together. One role alone is not a seam — and you will be asked to name which role your new package plays.
packages/shell/ is the canonical example. Consumers depend on the definition, never on a concrete provider, which is what keeps providers swappable.

There are 55 service keys in the system. Here are the ones you will meet first:

KeyCapabilityNotable providers
ctx.llmModel adaptersllm-deepseek, llm-pi-ai, llm-replay (tests)
ctx.sessionsThe session store and logcore — one implementation
ctx.toolsTool registry + guarded pipelinecore — one implementation
ctx.shellCommand executionbash-local, bash-sandbox, pwsh-*
ctx.fsFilesystem accessfs-local, fs-sandbox, fs-e2b
ctx.subprocessProcess spawningsubprocess-local, subprocess-e2b
ctx.sandboxConfinementsandbox-local (bwrap / Landlock / Seatbelt)
ctx.subagentsDelegation to child agents8 providers — see below
ctx.sessionPersistenceDurabilityJSONL, SQLite
ctx.approvalHuman confirmationper front end
The payoff, concretely

Filesystem and subprocess providers share one execution world. So pointing those two at a remote sandbox moves bash, the PTY terminals, and the language-server integration with them — no forks of any tool. The E2B remote-sandbox proof of concept is just three packages: e2b, fs-e2b, subprocess-e2b.

This is the test of whether a seam is real. When you design one, ask: "if someone swapped the provider, would everything downstream keep working?"

Chapter 10The rest of the system

You do not need these on day one, but you should know they exist and roughly where they live, so you recognise the names in review and in the event matrix.

SubsystemWhat it doesWhere
ScopePer-agent worlds. A tool or prompt section can be global or owned by exactly one agent; a scoped one shadows its global twin. tools.restrict() filters the global set per agent — and a filtered-away tool is absent from the prompt and refuses to execute.core/scope
PresetsCompose a whole agent from a preset cordis.yml, per session.preset/
SubagentsDelegation behind one interface. Providers include a fresh in-process child, a fork seeded from the parent's history, an out-of-process child over ACP or the SDK, a real Codex app-server, and a real Claude Code child via the official Agent SDK.subagent/
HooksBridges that read an existing Claude Code or Codex hooks.json and run those shell hooks faithfully, mapped onto this harness's own interception points. A "native hook" here is just an ordinary plugin.hooks/
CompactionShrinks history when the context fills — as a replace on the log surface, never a rewrite.compaction/
SpillOversized tool output goes to a file; the model gets a bounded preview plus a retrieval locator.spill/
JobsBackground work with job_list / job_output / job_kill control tools.jobs/
Session querySearch and trace across past sessions, including SQLite full-text search, exposed to the model as five session_* tools.session-query/
SkillsLoadable instruction packs the model can pull in on demand.skill/
Plan / Goal / SchedulePlan mode as logged state; durable same-session objectives; session-local scheduled follow-ups.plan/, goal/, schedule/
Workflow / RalphA workflow engine on worker threads, plus the "Ralph loop" — repeated fresh-agent attempts at one fixed objective.workflow/
Web GUI39 client packages registering into UI slots, plus the host half serving them. Typed RPC between the halves is generated by Typert.client/, host/, typert/
Self-modificationcordis_define / cordis_run / cordis_undefine: the model writes and mounts plugins into the runtime it is running inside, evaluated in a node:vm sandbox.extensions/
SDKs & protocolsJSON-RPC protocol + TypeScript client, an ACP automation server, MCP support, and a Python SDK that drives a bundled runtime binary.sdk/, acp/, mcp/, python/
InvariantsEvery package registers runtime checks under its own npm name. See Chapter 14.runtime-diagnostics/
Part III · The team

Chapter 11Who works on what

Read this before reading the tables

There is no CODEOWNERS file in this repository. What follows is derived from git history, so treat it as "who has been active where recently," not as a formal ownership map. Confirm with your lead before assuming someone is a reviewer.

Two caveats on the method. First, this clone is shallow: it covers roughly 2026-07-20 to 2026-08-14 — about four weeks and 863 non-merge commits. Longer-tenured ownership will not show up. Second, raw git log per directory is badly misleading here, because release commits touch 222 files across 52 areas at once and make one person look like they own everything. The numbers below exclude repo-wide sweeps (any commit touching more than six areas or sixty files), which removes 174 commits and leaves 689 focused ones.

By area — who to ask

AreaFocused commitsMost active, in order
Web UI (packages/client)176imccyu, _Kerman, Yichen Jiang, creatixchu, ZiyaZhang
Web app (apps/web)110Yichen Jiang, _Kerman, imccyu, creatixchu
Documentation (docs/)143Turtle, Yichen Jiang, Tianyi Cui, imccyu
Gates & generators (scripts/)123imccyu, Turtle, Yichen Jiang, Tianyi Cui
CLI (apps/cli)57Turtle, Yichen Jiang, Huanqi Cao, imccyu
Examples46Hypatia May, Yichen Jiang, pku-xht, imccyu
Web host (packages/host)37imccyu, _Kerman, Tianyi Cui, ZiyaZhang
Core spine (packages/core)30Chinesezjc, j-xiang, Yichen Jiang, imccyu
Bundles (packages/bundle)30Turtle, Tianyi Cui, imccyu, Huanqi Cao
Subagents29Hypatia May, pku-xht, j-xiang, imccyu
Boot glue (packages/boot)18Turtle, Huanqi Cao, Tianyi Cui
CI (.github/)18Chinesezjc, imccyu, Yichen Jiang
Python SDK14Yichen Jiang, imccyu, j-xiang, _Kerman
Agent presets14Yichen Jiang (dominant)
LLM adapters13j-xiang, Yichen Jiang, imccyu
Sandbox12Tianyi Cui, Huanqi Cao
API gateway (packages/api)12imccyu (dominant)
Self-modification (extensions/)12imccyu (dominant)
Feedback11Chinesezjc, ZiyaZhang, Turtle
Vendored Cordis11imccyu, Turtle
Typert (RPC codegen)6imccyu (sole)
MCP6Tianyi Cui (dominant)
Native Landlock addon4imccyu (sole)

By person — what each has been building

ContributorFocused commitsTheir patch of the map
imccyu122The broadest range in the repo. Web UI and host, the gate scripts, the API gateway, Typert, the self-modification toolset, the vendored framework, the native addon — and the release process.
Yichen Jiang108Also very broad, weighted to product surface: the web app and UI, documentation, agent presets, the Python SDK. Writes more Agent Notes than anyone.
Turtle77Documentation, the CLI, the gate scripts, bundles and boot glue. If you have a question about how the app assembles itself or where a doc belongs, this is the trail to follow.
Tianyi Cui70Cross-cutting standards work: gates, docs, the agent skills, plus sandbox, MCP, and schedule.
Chinesezjc60The core spine — the largest single contributor to packages/core in this window — plus CI, feedback, and code-runtime.
_Kerman57Front end. Client packages and the web app, with some host-side work.
creatixchu35Front end: client, web app, host, and the attachment subsystem.
ZiyaZhang34Front end plus the request-context packages and docs.
Huanqi Cao33Sandbox, boot, the CLI, workflow, and gate scripts. Systems-leaning.
pku-xht26Subagents, examples, and client work.
Hypatia May19Examples and subagents — the runnable compositions the snapshot tests boot.
j-xiang6Small but deep: core, subagents, LLM adapters, and Chinese documentation.
Yif, NI0317, fz10 / 7 / 5Front end — client packages and the web app.
xjt7Documentation and the bilingual translation corpus.

Two things the data tells you about how this team works

Everyone writes Agent Notes. .agents/notes/ is the single most-touched directory in the repository — 249 focused commits from 18 different people, ahead of every package and ahead of docs/. For nine of the top eleven contributors it is their #1 or #2 area. Design records are not bureaucracy layered on top of the work here; they are a large part of the work. Budget time for yours.

The front end is the busiest part of the product. packages/client plus apps/web is 286 focused commits, more than twice the core spine, docs, or scripts. If you are joining to work on "the agent," be aware that most day-to-day motion is in the browser half — and that the browser half is built out of the same plugin machinery as everything else.

Find your own reviewer

Before you open a pull request, run this on the files you changed. It beats any table:

git log --no-merges --format='%an' -- packages/<the-thing-you-touched> | sort | uniq -c | sort -rn | head
Part IV · Working here

Chapter 12Your first change

A good first task is adding a tool, because it exercises nearly every convention in the repo without requiring you to understand the loop deeply. The guided version is docs/user/develop/basic/tool.md; the contract reference is docs/cookbook/adding-a-tool.md.

Here is the full checklist of what a non-trivial change ships. Missing items are the most common reason a review stalls.

#DeliverableWhy it is required
1The package, named @deepseek-ai/dsh-<name>, in the right group, with a tsconfig.json referencing every workspace dependency.Naming and layout are gated mechanically.
2Unit tests under tests/ (never src/__tests__/), including an HMR-safety test — dispose the fiber, assert your contribution is gone.Per-file 100% line coverage is the CI gate. Disposal is what makes the plugin model real.
3A real-composition test — boot a test-only cordis.yml through the Loader and assert model-visible or durable output.Hand-built ctx.plugin(...) suites do not catch the "green tests, broken product" failure. See postmortem 0001.
4An ./invariant companion registered under your exact npm name.Every package has one. If you have nothing to check, export an empty installer whose comment starts No runtime invariant: and explains why — a gate rejects unexplained empties.
5A README with purpose, APIs, extension points, a Model Experience section, and ## Known Limitations and Deferred Work.All three sections are separately gated.
6An Agent Note in .agents/notes/, in the same PR.Required for anything beyond a mechanical edit. This is the decision record.
7A keyless snapshot scenario if the change is model-, protocol-, or human-visible.Package tests explicitly do not substitute for the assembled transcript.
8Docs updated in the same commit — affected READMEs and JSDoc."Docs accompany every code change" is a stated rule, not a nicety.

An ./invariant companion is smaller than it sounds. Here is a real one, from tool-todo, checking that every todo snapshot reaching the durable log is well-formed:

packages/todo/tool-todo/src/invariant.ts (abridged)
const PACKAGE_NAME = '@deepseek-ai/dsh-tool-todo'
export const name = 'tool-todo-invariant'
export const inject = ['invariants']

function validateTodos(value: unknown, fail: InvariantFailure): void {
  if (!Array.isArray(value)) fail('todo/write todos must be an array')
  const seen = new Set<string>()
  for (const item of value) {
    const { content, status } = item as Record<string, unknown>
    if (seen.has(content)) fail(`todo/write repeats content …`)
    if (!TODO_STATUSES.has(status)) fail(`todo/write carries unknown status …`)
  }
}

export const apply = (ctx: Context) =>
  Promise.resolve(ctx.invariants.register(PACKAGE_NAME, install))

Note what its JSDoc says it deliberately does not check — how many todos are in_progress — because that is per-deployment config, and "a log written while parallel work was allowed must still replay after a deployment tightens the policy." That is the level of care expected in these files.

Branching and pull requests

Chapter 13The gates

There are 124 scripts in scripts/ and a lot of them can fail your PR. The important thing to internalise first is which ones to run locally.

Do not run the full suite before every push

This is stated policy, not a shortcut: "Never default to the full suite or repeat a passing check for commit or push. CI owns exhaustive coverage and the platform matrix."

Match the evidence to what you changed: focused tests for behaviour, snapshots for model or user output, doc-sync for docs, build and hygiene for published paths, real-API e2e for provider behaviour. There is a skill — dsh-pre-push-checks — whose whole job is choosing the smallest sufficient set.

CommandWhat it covers
pnpm run testVitest unit tests.
pnpm run test:coverageThe CI coverage gate — per-file 100% lines on packages/*/*/src. Not test.
pnpm run test:snapshotKeyless replay of real example compositions, diffing transcripts and re-persisted logs. Filter with -t <name>.
pnpm run test:e2eReal-API tests. Self-skips without DEEPSEEK_API_KEY.
pnpm run test:webChromium browser snapshots. Required Linux PR gate.
pnpm run typecheck / linttsc / oxlint.
pnpm run doc-syncAll documentation gates at once — links, wrapping, budgets, translation pairing, catalog freshness, Mermaid, JSDoc.
pnpm run hygieneknip (dead exports) + publint + workspace constraints + NodeNext consumer check.
pnpm run duplicationCross-file clone detection.
pnpm run buildtsc emits lib/types; tsdown bundles the runtime.

Gates that surprise newcomers

Chapter 14Rules that will bite you

These are the conventions from AGENTS.md and packages/AGENTS.md that are non-obvious. Each one exists because something went wrong once.

Waterfall listeners must call next()

A waterfall is around-middleware. Returning without calling next() short-circuits the whole chain — every listener after you never runs. That is a legitimate design for a listener that owns a decision, and a bug for one that only annotates. Know which you are writing.

Use ctx.get(name) for optional services

Reserve ctx.<name> for services you declared in inject. The property proxy is topology-sensitive; ctx.get reads the global service store. Mixing them up produces "service is undefined" in one composition and not another.

Enforce a decision where it is made

Leaving a tool out of the prompt is not enforcement. Neither is a facade, a wrapper, or listener order — anything a direct caller can bypass. Test denial through the executor. The corollary is that a filtered-away tool must be indistinguishable from a nonexistent one.

No hardcoded tunables in plugins

Anything that varies by deployment is a validated Config field settable from cordis.yml. A DEFAULT_* constant or a test hook does not count as configurability. Protocol constants, external specifications, and security invariants stay fixed.

Explicit beats implicit at package boundaries

Defaulting is an explicit resolve(request): Spec step in the owning implementation, never a hidden ?? default buried in run(). The shell request/spec split is the template.

Misconfiguration fails loud

At load when self-contained, otherwise at the earliest resolvable point. Never silently skip a missing referent.

An empty catch must name what it swallows

…and why nothing else can reach it. Keep the try to a single statement.

Branded IDs, not bare strings

Opaque IDs that cross package boundaries use Branded<B>, so a SessionId cannot be passed where a CallId is expected. Structurally they are strings; at the type level they are not interchangeable.

Trust TypeScript at typed same-process boundaries

Do not add runtime validation or hostile-input tests for values the static interface already guarantees. Validate at the real boundaries: parsers, config, queues, model and tool JSON, files, workers, processes, and the wire. Defensive code everywhere is treated as noise here.

Write prose like a contract, not a transcript

This one catches almost everyone. Comments and docs state complete contracts and current behaviour — never reasoning narration. No "this used to…", no "a later PR will…", no restating what the code plainly does, no review history. Use direct, concrete terms; the standard explicitly asks you to write response fields rather than response shape, and to reserve the word contract for actual preconditions and guarantees. There is a skill, dsh-trim-cot-leakage, that exists purely to audit for this.

Why the prose rules are this strict

A large share of this codebase is written with agent assistance, and reasoning-transcript prose is the characteristic failure mode of agent-written text. The rules — and the gate — are a deliberate counterweight. It also means the comments you read in this repo are unusually trustworthy: when a JSDoc block says something is invariant, it is describing a real obligation, not a passing thought.

Chapter 15Where to look things up

Much of this repo's reference material is generated from source and verified fresh in CI. When you need a fact, prefer these over grepping.

QuestionGenerated answer
"What events exist, who fires them, who listens?"docs/event-producer-consumer.md
"What services exist and who implements them?"docs/capability-seams.md
"What config keys does this plugin accept?"docs/config-catalog.md
"What tools ship, with what schemas?"docs/tool-catalog.md
"What depends on what?"docs/module-graph.md
"What gets written to disk?"docs/persistence-catalog.md
"What does this word mean here?"docs/glossary.md
"How does one turn actually flow?"docs/agent-lifecycle.md · tool-execution-pipeline.md
"Why is it like this?".agents/notes/implemented/ — and .agents/notes/rejected/, which keeps 22 designs that were considered and turned down.

That last row is worth a habit. The expensive question in a mature codebase is never "how does this work" — it is "why isn't it the obvious other way." The rejected notes answer that directly, and reading a few of them will teach you this team's taste faster than anything else.

The five-minute version, if you remember nothing else
  1. Everything is a plugin, found by service key, registered as a reversible effect.
  2. The app is assembled from config rows, in layers, patchable by id.
  3. The session log is the source of truth. Model history is derived, compaction shadows rather than rewrites, and model-visible means logged.
  4. New behaviour attaches to an event, not to a branch in the loop.
  5. A change is not done until it has tests, an invariant companion, a README, an Agent Note, and a snapshot if anything visible changed.