Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Workflow Orchestration

How agents go beyond the basic loop to orchestrate multi-step and multi-agent work.

Four Philosophies

Qwen CodeGooseGrok BuildKimi Code
ParadigmTuring-complete JS workflow DSLDeclarative YAML session presetsHarness-driven state machineFlat subagent spawn/batch primitives
Orchestratornode:vm sandbox running user JSSame Agent::reply() loop as chatRust supervisor injecting synthetic turnsTS classes (SubagentHost/SubagentBatch)
New LLM calls per step?Yes — every agent() is a fresh sub-agentNo — recipe IS the one sessionMixed — executor stays in session; planner/verifier are spawnsYes — every subagent is a full nested Agent
Concurrency cap16 concurrent, 1000 total5 concurrent delegates1–5 skeptics (parallel), else sequentialRamp of 5 + 1/700ms
IsolationGit worktree per agent (opt-in)None beyond working_dirNone (single shared session)None — subagents share parent cwd
ResumeJournal replay keyed by call-sequence hashNone (restart from initial prompt)state.json persisted, manual /goal resumeresume(agentId) re-attaches
Budget modelExplicit token ceiling + double-check gateTurn-count onlyToken budget with monotonic ratchetNone inherited

Qwen Code — JavaScript Workflow DSL

Core files: repos/qwen-code/packages/core/src/tools/workflow/workflow.ts, agents/runtime/workflow-orchestrator.ts, workflow-sandbox.ts, workflow-budget.ts, workflow-journal.ts

Architecture

Four realms stacked inside one tool call:

Main session LLM  ──calls "workflow" tool──▶  WorkflowTool.execute()
                                                     │ allocates runId = wf_<16hex>
                                                     ▼
                                          WorkflowOrchestrator
                                          (shared ConcurrencyLimiter, budget, journal)
                                                     │ exposes agent()/parallel()/pipeline()/phase()/log()
                                                     ▼
                                          WorkflowSandbox (node:vm context)
                                          — runs user's JS script, NO filesystem/network
                                                     │ agent() calls bridge back to host
                                                     ▼
                                          AgentHeadless instances (real sub-agent LLM runs)

The script runs in a sandboxed node:vm context. All I/O happens through spawned agents — the script itself has no filesystem or network access.

Sandbox Security

  • Object.setPrototypeOf(bridge, null) severs prototype-chain escape
  • All values crossing the host↔vm boundary are JSON round-tripped (never passed by reference)
  • Math.random() and Date.now()/new Date() throw by design — determinism boundary so cached/replayed runs can’t diverge
  • Dual timeouts: V8 synchronous watchdog (30s) for infinite loops, plus async wall-clock cap (DEFAULT_MAX_WALL_CLOCK_MS = 30 * 60 * 1000)

Primitives

  • agent(prompt, opts?) — spawns AgentHeadless with max_turns: 50, max_time_minutes: 10. Nested Agent/Workflow/SendMessage/Monitor tools are disallowed in sub-agents.
  • parallel(thunks) — Promise.allSettled, rejected thunks become null (errors-as-data). Full-run abort is the sole exception.
  • pipeline(seed, ...stages) — threads each element through stages sequentially per item; null is the universal “drop” sentinel.
  • phase(title) — UI grouping for progress display.
  • log(msg) — capped at 10,000 lines; UI tail capped at 100.
  • workflow(nameOrRef, args?) — nested invocation, hard-limited to one level (child sandbox has no workflow implementation).

Concurrency

DEFAULT_MAX_AGENTS_PER_RUN = 1000;
HARD_MAX_AGENTS_PER_RUN_CEILING = 10_000;
HARD_MAX_CONCURRENCY_CEILING = 64;
// default: Math.max(1, Math.min(16, os.cpus().length - 2))

One ConcurrencyLimiter shared across the entire run — nested parallel()/pipeline() calls throttle against the same global cap.

Resume / Caching

Journal (workflow-journal.ts): <projectDir>/workflows/<runId>/journal.jsonl. Cache key per agent() call is a rolling hash: key = sha256(prefixHash + prompt + canonicalizeAgentOpts(opts)). Critical invariant: first cache miss invalidates every subsequent lookup (hadMiss flag), so the cache only ever serves a contiguous unbroken prefix.

Snapshots (workflow-snapshot.ts): whole-run summary for cross-restart /workflows history, capped at 30 retained.

Budget

Environment variable QWEN_CODE_MAX_TOKENS_PER_WORKFLOW, hard ceiling 100M. Enforcement is a double-check gate: once before dispatch (cheap) and once after a concurrency slot is acquired — because parallel() can queue many dispatches in one microtask burst, bounding worst-case overshoot to (concurrency_window − 1) × per_dispatch_tokens.

Worktree Isolation

provisionWorkflowWorktree fail-closed refuses: nesting worktrees when parent is already inside one, and provisioning from a dirty tree. Cleanup is fail-safe — never deletes a worktree with uncommitted/unmerged changes.

Error Handling

  • Stall watchdog (workflow-stall.ts): DEFAULT_STALL_MS = 60_000, MAX_STALL_ATTEMPTS = 3. Suspended while tool calls are in flight.
  • Sub-agent failure → rejected thunk, absorbed by parallel()/pipeline()’s errors-as-data contract.
  • Uncaught script errors preserved as WorkflowExecutionError with phases/logs/meta collected so far.

Goose — Recipe System (Declarative Session Presets)

Core files: repos/goose/crates/goose/src/recipe/mod.rs, template_recipe.rs, scheduler.rs, agents/schedule_tool.rs, platform_extensions/summon.rs

Key Insight: Recipe is Data, Not an Engine

A recipe is a declarative session preset — system prompt override, initial prompt, extension allowlist, model/temperature/max_turns, optional output schema, optional retry. All execution surfaces converge on the same Agent::reply() loop:

  • Interactive CLI recipe run: builder.rs stores recipe on session, constructs normal CliSession
  • Cron-triggered: scheduler.rs::execute_job builds fresh Agent::new() and calls agent.reply(...) headless
  • Delegated subagent: subagent_handler.rs::get_agent_messages builds Agent::with_config

YAML Format

#![allow(unused)]
fn main() {
pub struct Recipe {
    pub version, title, description,
    pub instructions: Option<String>,
    pub prompt: Option<String>,
    pub extensions: Option<Vec<ExtensionConfig>>,
    pub settings: Option<Settings>,        // model, temperature, max_turns
    pub parameters: Option<Vec<RecipeParameter>>,
    pub response: Option<Response>,        // json_schema validated
    pub sub_recipes: Option<Vec<SubRecipe>>,
    pub retry: Option<RetryConfig>,
}
}

Templating

MiniJinja, two-phase: a lenient discovery pass to find declared variables, then a strict render pass (UndefinedBehavior::Strict) — undefined variable = hard error. A preprocessing step wraps literal {{...}} prose in {% raw %} blocks.

Scheduled Execution

Real in-process async cron via tokio-cron-scheduler (not polling, not OS daemon). Security hardening:

  • MAX_SCHEDULE_RECIPE_BYTES = 1MB
  • O_NONBLOCK|O_NOFOLLOW opens rejecting FIFOs and symlink swaps
  • Recipes copied to scheduled_recipes/{job_id}.{ext} at 0o600
  • On restart: clears stale currently_running flags (crash cleanup, not resume)

Multi-Agent: Two Orthogonal Extensions

summon — delegate/load tools:

  • Hard-caps delegation depth at 1 (subagent’s delegate calls rejected)
  • Coaches model on parallel-vs-sequential use and file-partitioning for concurrent writers
  • GOOSE_MAX_BACKGROUND_TASKS default 5

orchestrator — manages long-lived peer sessions:

  • list_sessions, start_agent, send_message, interrupt_agent
  • Peer model, not parent-child

Sub-Recipes

Executed through summon: Recipe::ensure_summon_for_subrecipes auto-injects the extension. SubRecipe.sequential_when_repeated is declared but never consumed anywhere in the execution path — parallelism vs sequencing is left entirely to the LLM’s tool-calling behavior.

Retry

RetryConfig{max_retries, checks: Vec<SuccessCheck::Shell>, on_failure, timeout_seconds}. On failure, reset_status_for_retry wipes conversation back to initial_messages — each retry is a hard restart.

Gap found: Retry is wired for CLI and summon-delegated runs, but scheduler.rs::execute_job builds SessionConfig{..., retry_config: None} — retry is silently disabled for cron-triggered runs even if the recipe declares it.


Grok Build — Goal Tracking Pipeline (Harness-Driven State Machine)

Core files: repos/grok-build/crates/codegen/xai-grok-shell/src/session/goal_tracker.rs, goal_orchestrator.rs, goal_classifier.rs, goal_planner.rs, goal_strategist.rs, goal_summarizer.rs

Architecture: Hybrid

The executor/implementer work happens in the same parent session, advanced by the harness injecting synthetic “user” turns. But planner, verifier, strategist, and summarizer are genuine separate subagent spawns dispatched from harness Rust code — never via a model-visible tool call, keeping the parent’s transcript clean.

Parent ACP session (executor)
   │  harness injects continuation directive each round
   │
   ├─▶ Planner subagent   (once, at goal creation, writes plan.md)
   ├─▶ Evaluator (cheap model, every round) → Continue | CandidateComplete | Blocked
   ├─▶ Verifier panel (N adversarial skeptics, on CandidateComplete)
   ├─▶ Strategist subagent (after repeated stalls)
   └─▶ Summarizer subagent (on Achieved)

State Machine

#![allow(unused)]
fn main() {
GoalPhase { Idle, Planning, Executing }
GoalStatus { Active, UserPaused, BackOffPaused, NoProgressPaused,
             InfraPaused, Blocked, BudgetLimited, Complete }
}

Safety: custom Deserialize maps any unknown wire value to UserPaused — “so a corrupt or forward-version snapshot can never resurrect as a self-driving goal.” On restart, from_snapshot resets Planning/Executing → Idle and Active → UserPaused; the user must explicitly /goal resume.

Adversarial Verification

goal_classifier.rs spawns GOAL_VERIFIER_SKEPTIC_COUNT (default 3, clamped 1–5) adversarial subagents with real tool access (read/search/list/run-command):

  • Escalating panel: skeptic 0 runs alone first; refuted && confidence==High from skeptic 0 alone is decisive (short-circuits remaining spawns)
  • Majority-vote aggregation for the rest
  • Outcome: Achieved | NotAchieved{gaps_summary, gap_fingerprint} | Blocked | FailOpenAchieved
  • Verifier prompt: “Default to refuted: true if uncertain”
  • Anti-ratchet section: “converge, don’t re-litigate”

Control Flow Per Round

  1. Cheap/fast evaluator reads bounded transcript (TRANSCRIPT_MAX_BYTES = 32KB) → Continue | CandidateComplete | Blocked
  2. Token-budget check; ends goal if exceeded
  3. CandidateComplete → verifier panel. Blocked (3× same key) → auto-pause. Continue → next continuation directive from plan’s first unchecked item
  4. Verifier NotAchieved → stall counter; 2 identical gap_fingerprints → auto-pause; strategist fires every N failures

Error Handling: Fail-Open/Fail-Closed Split

  • Verifier/evaluator infra failures → fail open (never silently block user)
  • Planner failure → fail closed (pauses goal)
  • Strategist/summarizer → fail open (best-effort)
  • Two independent “3-strike” rules gate auto-pausing

Persistence

<session_dir>/goal/state.json — pretty-JSON serialization of GoalOrchestration, written atomically (temp+rename) on every state-changing event.

Budget

Optional user-specified token cap (setup_goal(objective, token_budget)), enforced every round. Uses “monotonic high-water-mark, positive-delta-only” accounting so context compaction never makes cumulative usage appear to shrink.


Kimi Code — Flat Subagent Spawn/Batch

Core files: repos/kimi-code/packages/agent-core/src/session/subagent-host.ts, subagent-batch.ts

Architecture: Same Loop, Nested

SessionSubagentHost.spawn() creates the same Agent class the main loop uses, sharing the parent’s model-generate function and cwd, running its own full turn to completion. No separate orchestration engine.

Batch Scheduler

SubagentBatch<T> — hand-rolled, not Promise.all:

Normal phase: INITIAL_LAUNCH_LIMIT = 5 immediately, then +1 every 700ms. Optional hard cap via KIMI_CODE_AGENT_SWARM_MAX_CONCURRENCY.

Rate-limit phase: On any provider rate-limit error:

  • Requeue task at front with exponential backoff (RATE_LIMIT_RETRY_BASE_MS = 3000, factor 2)
  • Shrink capacity by 1 (min 1, throttled to once per 2s)
  • Recover +1 after 3 minutes clean

Results preserve input order regardless of completion order. Cancellation preserves partial results with distinguishable started/not_started states.

No Isolation

Subagents run against the identical working tree as their parent. No worktree, no filesystem sandbox.

Explicitly Absent

  • No workflow DSL, no pipeline/parallel primitives
  • No declarative recipes
  • No goal state machine for multi-step orchestration
  • GoalMode exists but is a flat, single-agent budget/lifecycle tracker — not multi-step

Cron (Orthogonal)

CronManager wraps a CronScheduler (1s poll or SIGUSR1 tick), persists tasks to <sessionDir>/cron/<id>.json. Explicitly disabled for subagents (test proves ctx.agent.cron is null for type:'sub').

Error Handling

  • Timeout is recoverable: timed-out subagent’s context preserved, caller gets resume_hint
  • Mid-task errors classified by terminal reason: PROVIDER_RATE_LIMIT → batch requeue, others → generic error result
  • Cancellation only recurses into foreground children

Synthesis: What Each System Is Really Solving

Qwen Code treats orchestration as a programming problem — give the author a real, sandboxed, deterministic scripting language with fan-out/fan-in, caching, and hard resource ceilings. Most general, most heavily engineered. The only one with resumable execution cache, security-hardened sandbox, and worktree isolation baked into the orchestration primitive.

Goose treats orchestration as a configuration problem — recipes don’t add execution semantics, they parameterize the same single-agent loop. Most interesting engineering is scheduler security (TOCTOU/symlink/FIFO defenses) and the deliberate separation of “recipe” (data) from “summon”/“orchestrator” (multi-agent mechanisms).

Grok Build treats orchestration as a verification problem — the interesting design is not fan-out breadth but adversarial rigor. A self-skeptical panel that defaults to disbelief, an explicit fail-open/fail-closed split by role, and a state machine engineered so a corrupted snapshot can never resurrect an unattended self-driving loop.

Kimi Code deliberately does not solve orchestration beyond spawn-and-batch. Its engineering effort goes into making the batch scheduler robust against real-world provider rate limiting (hand-rolled ramp + backoff + capacity-recovery) rather than higher-level control flow.