Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Error Recovery & Doom Loops

How agents detect they’re stuck, enforce limits, and recover from repeated failures.

The Core Finding

No agent has genuine semantic “doom loop” detection (noticing the same failing action repeated). What exists are proxies — turn/step ceilings, token/compaction budgets, and narrow circuit breakers. The model itself is the primary recovery mechanism.

Cross-Agent Comparison

AgentHard Turn/Step CapLoop DetectionRecovery StrategyModel Informed?
Goose1000 turns (configurable)RepetitionInspector (disabled by default); stop-hook denial cap (8)Stop-hook forces continuation; retry resets conversationYes
Grok BuildPer-goal token budgetStall counter (2× identical gap_fingerprint → pause); evaluator 3× Blocked → pauseStrategist subagent for course correction; auto-pauseYes (continuation directive)
Qwen CodeWorkflow: 1000 agents, wall-clock 30minStall watchdog (60s inactivity, 3 attempts)Stall → abort workflow; step limits per sub-agentYes (workflow error)
CodexNone (optional token budget only)Guardian denial breaker (3 consecutive / 10-of-50)None generic — relies on model + compactionNo
OpenCodesteps config (default Infinity)None (explicit TODO in source)Forced text-only turn at limitOnly if limit configured
Kilocode25 steps hard capCompaction-attempt cap (3); ConsecutiveMistakeError scaffolded but unusedStepLimitExceededError terminates runYes
PiConfigurable max stepsNoneError fed back to model as tool resultNo

Goose — Turn Limits + Stop-Hook Denial Cap

Core file: repos/goose/crates/goose/src/agents/agent.rs

Turn Limits

#![allow(unused)]
fn main() {
const DEFAULT_MAX_TURNS: u32 = 1000;
const DEFAULT_STOP_HOOK_BLOCK_CAP: u32 = 8;
const MAX_EMPTY_TURN_RETRIES: u32 = 3;
}

Hard break at max_turns — configurable per-session or via GOOSE_MAX_TURNS env var. Fires regardless of stop hooks. The model receives MAX_TURNS_MESSAGE: “I’ve reached the maximum number of actions I can do without user input.”

RepetitionInspector (Disabled by Default)

repos/goose/crates/goose/src/tool_monitor.rs:

#![allow(unused)]
fn main() {
pub fn check_tool_call(&mut self, tool_call: CallToolRequestParams) -> bool {
    if last.matches(&internal_call) {
        self.repeat_count += 1;
        if self.repeat_count > self.max_repetitions.unwrap() { return false; }
    } else { self.repeat_count = 1; }
}
}

When triggered, denies the tool call via ToolInspectionManager rather than aborting. Not exercised by default since no repetition limit is configured out of the box.

Stop-Hook Denial System

Hooks can deny the agent from stopping (HookDecision::Deny). On denial:

  1. Injects invisible synthetic user message telling model to address the denial
  2. Loops again (forced continuation)
  3. Cap: DEFAULT_STOP_HOOK_BLOCK_CAP = 8 consecutive denials → force-stop

Key interaction: stop-hook-denial retries do NOT consume turn budget (increment skipped), so a policy plugin that keeps denying is bounded only by the 8-denial cap.

Transport Retry

repos/goose/crates/goose-provider-types/src/retry.rs: exponential backoff with jitter (initial 1s, multiplier 2×, max 30s, 3 retries). Only retries RateLimitExceeded | ServerError | NetworkError.

Task-Level Retry

repos/goose/crates/goose/src/agents/retry.rs: RetryManager runs SuccessCheck::Shell checks. On failure, wipes conversation back to initial messages and retries from scratch. Max retries configurable per recipe.


Grok Build — Adversarial Stall Detection

Core files: repos/grok-build/crates/codegen/xai-grok-shell/src/session/goal_tracker.rs, goal_orchestrator.rs, goal_classifier.rs

Stall Detection (Goal System)

Per round, after the verifier panel returns NotAchieved:

  • Increments a stall counter
  • Compares gap_fingerprint with previous round
  • 2 identical fingerprints in a row → auto-pause as stalled
  • Relaxed to 5 while a strategist restructure is active

Evaluator-Based Pause

Cheap/fast evaluator model runs every round → Continue | CandidateComplete | Blocked:

  • 3 consecutive Blocked decisions on the same key → auto-pause
  • classifier_max_runs (default 10) caps total verification attempts

Strategist (Course Correction)

Fires after N consecutive verification failures. A separate subagent that recommends structural changes to the approach, buying a +3-run cap bonus before the next auto-pause.

Circuit Breaker for Permission Denials

AUTO_DENY_CONSECUTIVE_LIMIT = 3, AUTO_DENY_TOTAL_LIMIT = 20 — escalates repeated auto-mode denials back to user: “Take a safer approach… do not retry this exact action.”

Budget Enforcement

Per-goal token budget with “monotonic high-water-mark, positive-delta-only” accounting — context compaction never makes cumulative usage appear to shrink.


Qwen Code — Workflow Stall Watchdog

Core file: repos/qwen-code/packages/core/src/agents/runtime/workflow-stall.ts

Stall Detection

DEFAULT_STALL_MS = 60_000
MAX_STALL_ATTEMPTS = 3
  • Suspended while any tool call is in flight (slow shell commands don’t trigger)
  • Not armed until the first activity event (time-to-first-response doesn’t count)
  • 3 stall timeouts → abort the workflow

Sub-Agent Limits

Per agent() call: max_turns: 50, max_time_minutes: 10. Failure becomes a rejected thunk absorbed by parallel()/pipeline()’s errors-as-data contract.

Workflow-Level Limits

  • 1000 total agents per run (hard cap, call 1001 throws)
  • Wall-clock timeout: 30 minutes default
  • Token budget: double-check gate (before dispatch + after slot acquired)

Codex — Minimal: Token Budget Only

Core file: repos/codex/codex-rs/core/src/session/turn.rs

No Turn/Step Limits

No max_turns, max_steps, or step-count cap anywhere. The loop runs until the model produces a final message, an unretryable error occurs, or the optional token budget is exhausted.

Token Budget (Optional)

RolloutBudgetConfig with limit_tokens: i64. Records weighted token usage (output × sampling weight + non-cached input × prefill weight). When exceeded → SessionBudgetExceeded error terminates the turn.

Guardian Rejection Circuit Breaker

#![allow(unused)]
fn main() {
pub const MAX_CONSECUTIVE_GUARDIAN_DENIALS_PER_TURN: u32 = 3;
pub const MAX_RECENT_AUTO_REVIEW_DENIALS_PER_TURN: u32 = 10;
pub const AUTO_REVIEW_DENIAL_WINDOW_SIZE: usize = 50;
}

Scoped to the auto-approval reviewer, not general failures. Fires when reviewer rejects 3 consecutive or 10-of-last-50, aborting the turn.

Compaction as Implicit Loop Prevention

Code comment (turn.rs:394):

#![allow(unused)]
fn main() {
// as long as compaction works well in getting us way below the token limit,
// we shouldn't worry about being in an infinite loop.
}

Explicit acknowledgment: compaction is the de facto soft mechanism preventing infinite loops from hitting a hard ceiling.


OpenCode — Explicit Gap (TODO in Source)

Core file: repos/opencode/packages/core/src/session/runner/llm.ts

Step Limit (Config-Driven, Default Infinity)

const isLastStep = agent.info?.steps !== undefined && currentStep >= agent.info.steps

When hit, tools are omitted and forced text-only turn appended via MAX_STEPS_PROMPT:

“CRITICAL - MAXIMUM STEPS REACHED. The maximum number of steps allowed for this task has been reached. Tools are disabled until next user input.”

No Loop Detection

Source acknowledges the gap (runner/llm.ts:55):

// [ ] Bound provider retries and repeated identical tool calls.

Provider Retry

MAX_RETRIES = 2, BASE_DELAY_MS = 500, MAX_DELAY_MS = 10_000. Only for retryable HTTP codes (429/503/504/529).


Kilocode — Hard Cap + Compaction Guard

Core file: repos/kilocode/packages/core/src/session/runner/llm.ts

Hard Step Cap (Added Over OpenCode)

const MAX_STEPS = 25
for (let step = 0; step < MAX_STEPS; step++) { ... }
if (needsContinuation)
    return yield* new StepLimitExceededError({ sessionID, limit: MAX_STEPS })

StepLimitExceededError is a typed error that fails the whole run.

Compaction-Attempt Guard

export const MAX_COMPACTION_ATTEMPTS = 3

Prevents infinite compaction loops. Comment: // kilocode_change - cap compaction attempts per turn to avoid infinite loops. OpenCode has no equivalent.

ConsecutiveMistakeError (Scaffolded, Not Live)

export type ConsecutiveMistakeReason = "no_tools_used" | "tool_repetition" | "unknown"

Defined as telemetry scaffolding but no live call site constructs this error — appears to be planned but not yet integrated.


Design Patterns

1. Proxies, Not Detection

No agent detects “the model is doing the same failing thing repeatedly” at a semantic level. Instead they use:

  • Turn/step ceilings (Goose 1000, Kilocode 25, configurable elsewhere)
  • Token budgets (Codex, Grok Build per-goal)
  • Wall-clock timeouts (Qwen Code workflows 30min)
  • Compaction as soft cap (Codex explicitly, others implicitly)

2. Model-as-Recovery-Agent

The most common “recovery” is simply feeding the error back to the model as a tool result and trusting it to adapt. This is the only strategy in Codex, Pi, and OpenCode.

3. Conversation Reset (Nuclear Option)

Both Goose (task-level retry) and Grok Build (on stall + strategist failure) can wipe the conversation back to initial messages and start fresh. This is the most aggressive recovery — it discards all work done so far.

4. Escalation to User

  • Grok Build: auto-pause on stall, requires /goal resume
  • Goose: MAX_TURNS_MESSAGE asks user to intervene
  • Grok Build permission system: circuit breaker after 3/20 denials

5. The Gap is Acknowledged

Both Codex (code comment about compaction) and OpenCode (explicit TODO) acknowledge that proper doom-loop detection doesn’t exist. Kilocode’s ConsecutiveMistakeError scaffolding shows intent to address it. This is a known unsolved problem across the field.