Workflow Orchestration
How agents go beyond the basic loop to orchestrate multi-step and multi-agent work.
Four Philosophies
| Qwen Code | Goose | Grok Build | Kimi Code | |
|---|---|---|---|---|
| Paradigm | Turing-complete JS workflow DSL | Declarative YAML session presets | Harness-driven state machine | Flat subagent spawn/batch primitives |
| Orchestrator | node:vm sandbox running user JS | Same Agent::reply() loop as chat | Rust supervisor injecting synthetic turns | TS classes (SubagentHost/SubagentBatch) |
| New LLM calls per step? | Yes — every agent() is a fresh sub-agent | No — recipe IS the one session | Mixed — executor stays in session; planner/verifier are spawns | Yes — every subagent is a full nested Agent |
| Concurrency cap | 16 concurrent, 1000 total | 5 concurrent delegates | 1–5 skeptics (parallel), else sequential | Ramp of 5 + 1/700ms |
| Isolation | Git worktree per agent (opt-in) | None beyond working_dir | None (single shared session) | None — subagents share parent cwd |
| Resume | Journal replay keyed by call-sequence hash | None (restart from initial prompt) | state.json persisted, manual /goal resume | resume(agentId) re-attaches |
| Budget model | Explicit token ceiling + double-check gate | Turn-count only | Token budget with monotonic ratchet | None inherited |
Qwen Code — JavaScript Workflow DSL
Core files: repos/qwen-code/packages/core/src/tools/workflow/workflow.ts, agents/runtime/workflow-orchestrator.ts, workflow-sandbox.ts, workflow-budget.ts, workflow-journal.ts
Architecture
Four realms stacked inside one tool call:
Main session LLM ──calls "workflow" tool──▶ WorkflowTool.execute()
│ allocates runId = wf_<16hex>
▼
WorkflowOrchestrator
(shared ConcurrencyLimiter, budget, journal)
│ exposes agent()/parallel()/pipeline()/phase()/log()
▼
WorkflowSandbox (node:vm context)
— runs user's JS script, NO filesystem/network
│ agent() calls bridge back to host
▼
AgentHeadless instances (real sub-agent LLM runs)
The script runs in a sandboxed node:vm context. All I/O happens through spawned agents — the script itself has no filesystem or network access.
Sandbox Security
Object.setPrototypeOf(bridge, null)severs prototype-chain escape- All values crossing the host↔vm boundary are JSON round-tripped (never passed by reference)
Math.random()andDate.now()/new Date()throw by design — determinism boundary so cached/replayed runs can’t diverge- Dual timeouts: V8 synchronous watchdog (30s) for infinite loops, plus async wall-clock cap (
DEFAULT_MAX_WALL_CLOCK_MS = 30 * 60 * 1000)
Primitives
agent(prompt, opts?)— spawnsAgentHeadlesswithmax_turns: 50,max_time_minutes: 10. NestedAgent/Workflow/SendMessage/Monitortools are disallowed in sub-agents.parallel(thunks)—Promise.allSettled, rejected thunks becomenull(errors-as-data). Full-run abort is the sole exception.pipeline(seed, ...stages)— threads each element through stages sequentially per item;nullis the universal “drop” sentinel.phase(title)— UI grouping for progress display.log(msg)— capped at 10,000 lines; UI tail capped at 100.workflow(nameOrRef, args?)— nested invocation, hard-limited to one level (child sandbox has noworkflowimplementation).
Concurrency
DEFAULT_MAX_AGENTS_PER_RUN = 1000;
HARD_MAX_AGENTS_PER_RUN_CEILING = 10_000;
HARD_MAX_CONCURRENCY_CEILING = 64;
// default: Math.max(1, Math.min(16, os.cpus().length - 2))
One ConcurrencyLimiter shared across the entire run — nested parallel()/pipeline() calls throttle against the same global cap.
Resume / Caching
Journal (workflow-journal.ts): <projectDir>/workflows/<runId>/journal.jsonl. Cache key per agent() call is a rolling hash: key = sha256(prefixHash + prompt + canonicalizeAgentOpts(opts)). Critical invariant: first cache miss invalidates every subsequent lookup (hadMiss flag), so the cache only ever serves a contiguous unbroken prefix.
Snapshots (workflow-snapshot.ts): whole-run summary for cross-restart /workflows history, capped at 30 retained.
Budget
Environment variable QWEN_CODE_MAX_TOKENS_PER_WORKFLOW, hard ceiling 100M. Enforcement is a double-check gate: once before dispatch (cheap) and once after a concurrency slot is acquired — because parallel() can queue many dispatches in one microtask burst, bounding worst-case overshoot to (concurrency_window − 1) × per_dispatch_tokens.
Worktree Isolation
provisionWorkflowWorktree fail-closed refuses: nesting worktrees when parent is already inside one, and provisioning from a dirty tree. Cleanup is fail-safe — never deletes a worktree with uncommitted/unmerged changes.
Error Handling
- Stall watchdog (
workflow-stall.ts):DEFAULT_STALL_MS = 60_000,MAX_STALL_ATTEMPTS = 3. Suspended while tool calls are in flight. - Sub-agent failure → rejected thunk, absorbed by
parallel()/pipeline()’s errors-as-data contract. - Uncaught script errors preserved as
WorkflowExecutionErrorwith phases/logs/meta collected so far.
Goose — Recipe System (Declarative Session Presets)
Core files: repos/goose/crates/goose/src/recipe/mod.rs, template_recipe.rs, scheduler.rs, agents/schedule_tool.rs, platform_extensions/summon.rs
Key Insight: Recipe is Data, Not an Engine
A recipe is a declarative session preset — system prompt override, initial prompt, extension allowlist, model/temperature/max_turns, optional output schema, optional retry. All execution surfaces converge on the same Agent::reply() loop:
- Interactive CLI recipe run:
builder.rsstores recipe on session, constructs normalCliSession - Cron-triggered:
scheduler.rs::execute_jobbuilds freshAgent::new()and callsagent.reply(...)headless - Delegated subagent:
subagent_handler.rs::get_agent_messagesbuildsAgent::with_config
YAML Format
#![allow(unused)]
fn main() {
pub struct Recipe {
pub version, title, description,
pub instructions: Option<String>,
pub prompt: Option<String>,
pub extensions: Option<Vec<ExtensionConfig>>,
pub settings: Option<Settings>, // model, temperature, max_turns
pub parameters: Option<Vec<RecipeParameter>>,
pub response: Option<Response>, // json_schema validated
pub sub_recipes: Option<Vec<SubRecipe>>,
pub retry: Option<RetryConfig>,
}
}
Templating
MiniJinja, two-phase: a lenient discovery pass to find declared variables, then a strict render pass (UndefinedBehavior::Strict) — undefined variable = hard error. A preprocessing step wraps literal {{...}} prose in {% raw %} blocks.
Scheduled Execution
Real in-process async cron via tokio-cron-scheduler (not polling, not OS daemon). Security hardening:
MAX_SCHEDULE_RECIPE_BYTES = 1MBO_NONBLOCK|O_NOFOLLOWopens rejecting FIFOs and symlink swaps- Recipes copied to
scheduled_recipes/{job_id}.{ext}at0o600 - On restart: clears stale
currently_runningflags (crash cleanup, not resume)
Multi-Agent: Two Orthogonal Extensions
summon — delegate/load tools:
- Hard-caps delegation depth at 1 (subagent’s delegate calls rejected)
- Coaches model on parallel-vs-sequential use and file-partitioning for concurrent writers
GOOSE_MAX_BACKGROUND_TASKSdefault 5
orchestrator — manages long-lived peer sessions:
list_sessions,start_agent,send_message,interrupt_agent- Peer model, not parent-child
Sub-Recipes
Executed through summon: Recipe::ensure_summon_for_subrecipes auto-injects the extension. SubRecipe.sequential_when_repeated is declared but never consumed anywhere in the execution path — parallelism vs sequencing is left entirely to the LLM’s tool-calling behavior.
Retry
RetryConfig{max_retries, checks: Vec<SuccessCheck::Shell>, on_failure, timeout_seconds}. On failure, reset_status_for_retry wipes conversation back to initial_messages — each retry is a hard restart.
Gap found: Retry is wired for CLI and summon-delegated runs, but scheduler.rs::execute_job builds SessionConfig{..., retry_config: None} — retry is silently disabled for cron-triggered runs even if the recipe declares it.
Grok Build — Goal Tracking Pipeline (Harness-Driven State Machine)
Core files: repos/grok-build/crates/codegen/xai-grok-shell/src/session/goal_tracker.rs, goal_orchestrator.rs, goal_classifier.rs, goal_planner.rs, goal_strategist.rs, goal_summarizer.rs
Architecture: Hybrid
The executor/implementer work happens in the same parent session, advanced by the harness injecting synthetic “user” turns. But planner, verifier, strategist, and summarizer are genuine separate subagent spawns dispatched from harness Rust code — never via a model-visible tool call, keeping the parent’s transcript clean.
Parent ACP session (executor)
│ harness injects continuation directive each round
│
├─▶ Planner subagent (once, at goal creation, writes plan.md)
├─▶ Evaluator (cheap model, every round) → Continue | CandidateComplete | Blocked
├─▶ Verifier panel (N adversarial skeptics, on CandidateComplete)
├─▶ Strategist subagent (after repeated stalls)
└─▶ Summarizer subagent (on Achieved)
State Machine
#![allow(unused)]
fn main() {
GoalPhase { Idle, Planning, Executing }
GoalStatus { Active, UserPaused, BackOffPaused, NoProgressPaused,
InfraPaused, Blocked, BudgetLimited, Complete }
}
Safety: custom Deserialize maps any unknown wire value to UserPaused — “so a corrupt or forward-version snapshot can never resurrect as a self-driving goal.” On restart, from_snapshot resets Planning/Executing → Idle and Active → UserPaused; the user must explicitly /goal resume.
Adversarial Verification
goal_classifier.rs spawns GOAL_VERIFIER_SKEPTIC_COUNT (default 3, clamped 1–5) adversarial subagents with real tool access (read/search/list/run-command):
- Escalating panel: skeptic 0 runs alone first;
refuted && confidence==Highfrom skeptic 0 alone is decisive (short-circuits remaining spawns) - Majority-vote aggregation for the rest
- Outcome:
Achieved | NotAchieved{gaps_summary, gap_fingerprint} | Blocked | FailOpenAchieved - Verifier prompt: “Default to
refuted: trueif uncertain” - Anti-ratchet section: “converge, don’t re-litigate”
Control Flow Per Round
- Cheap/fast evaluator reads bounded transcript (
TRANSCRIPT_MAX_BYTES = 32KB) →Continue | CandidateComplete | Blocked - Token-budget check; ends goal if exceeded
CandidateComplete→ verifier panel.Blocked(3× same key) → auto-pause.Continue→ next continuation directive from plan’s first unchecked item- Verifier
NotAchieved→ stall counter; 2 identicalgap_fingerprints → auto-pause; strategist fires every N failures
Error Handling: Fail-Open/Fail-Closed Split
- Verifier/evaluator infra failures → fail open (never silently block user)
- Planner failure → fail closed (pauses goal)
- Strategist/summarizer → fail open (best-effort)
- Two independent “3-strike” rules gate auto-pausing
Persistence
<session_dir>/goal/state.json — pretty-JSON serialization of GoalOrchestration, written atomically (temp+rename) on every state-changing event.
Budget
Optional user-specified token cap (setup_goal(objective, token_budget)), enforced every round. Uses “monotonic high-water-mark, positive-delta-only” accounting so context compaction never makes cumulative usage appear to shrink.
Kimi Code — Flat Subagent Spawn/Batch
Core files: repos/kimi-code/packages/agent-core/src/session/subagent-host.ts, subagent-batch.ts
Architecture: Same Loop, Nested
SessionSubagentHost.spawn() creates the same Agent class the main loop uses, sharing the parent’s model-generate function and cwd, running its own full turn to completion. No separate orchestration engine.
Batch Scheduler
SubagentBatch<T> — hand-rolled, not Promise.all:
Normal phase: INITIAL_LAUNCH_LIMIT = 5 immediately, then +1 every 700ms. Optional hard cap via KIMI_CODE_AGENT_SWARM_MAX_CONCURRENCY.
Rate-limit phase: On any provider rate-limit error:
- Requeue task at front with exponential backoff (
RATE_LIMIT_RETRY_BASE_MS = 3000, factor 2) - Shrink capacity by 1 (min 1, throttled to once per 2s)
- Recover +1 after 3 minutes clean
Results preserve input order regardless of completion order. Cancellation preserves partial results with distinguishable started/not_started states.
No Isolation
Subagents run against the identical working tree as their parent. No worktree, no filesystem sandbox.
Explicitly Absent
- No workflow DSL, no pipeline/parallel primitives
- No declarative recipes
- No goal state machine for multi-step orchestration
GoalModeexists but is a flat, single-agent budget/lifecycle tracker — not multi-step
Cron (Orthogonal)
CronManager wraps a CronScheduler (1s poll or SIGUSR1 tick), persists tasks to <sessionDir>/cron/<id>.json. Explicitly disabled for subagents (test proves ctx.agent.cron is null for type:'sub').
Error Handling
- Timeout is recoverable: timed-out subagent’s context preserved, caller gets
resume_hint - Mid-task errors classified by terminal reason:
PROVIDER_RATE_LIMIT→ batch requeue, others → generic error result - Cancellation only recurses into foreground children
Synthesis: What Each System Is Really Solving
Qwen Code treats orchestration as a programming problem — give the author a real, sandboxed, deterministic scripting language with fan-out/fan-in, caching, and hard resource ceilings. Most general, most heavily engineered. The only one with resumable execution cache, security-hardened sandbox, and worktree isolation baked into the orchestration primitive.
Goose treats orchestration as a configuration problem — recipes don’t add execution semantics, they parameterize the same single-agent loop. Most interesting engineering is scheduler security (TOCTOU/symlink/FIFO defenses) and the deliberate separation of “recipe” (data) from “summon”/“orchestrator” (multi-agent mechanisms).
Grok Build treats orchestration as a verification problem — the interesting design is not fan-out breadth but adversarial rigor. A self-skeptical panel that defaults to disbelief, an explicit fail-open/fail-closed split by role, and a state machine engineered so a corrupted snapshot can never resurrect an unattended self-driving loop.
Kimi Code deliberately does not solve orchestration beyond spawn-and-batch. Its engineering effort goes into making the batch scheduler robust against real-world provider rate limiting (hand-rolled ramp + backoff + capacity-recovery) rather than higher-level control flow.