Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Hashline Anchor Schemes

A deep dive into Grok Build’s unique anchor-based file editing system — the only non-exact-match, non-diff edit strategy in the agents studied.

File Map

Base directory: repos/grok-build/crates/codegen/xai-grok-tools/src/implementations/grok_build_hashline/

FileRole
scheme.rsThree AnchorScheme implementations, Anchor/ParsedAnchor types, find_shifted recovery
anchor.rsRe-exports + split_lines, generate_for_content, validate_against_content helpers
../../util/hash.rsFNV-1a hashing primitives and letter encoding
config.rsHashlineSchemeParams (per-session config), build_scheme()
read_file.rshashline_read tool; produces LINE:LOCAL[:CONTEXT]→CONTENT output
edit/{mod.rs,apply.rs,types.rs}hashline_edit tool; anchor validation, shift recovery, batch apply
grep.rshashline_grep; injects anchors into ripgrep output
benchmark.rsOffline microbenchmark comparing all three schemes

Hash Primitives

FNV-1a (util/hash.rs)

Standard 32-bit FNV-1a: offset basis 2_166_136_261, prime 16_777_619.

Line Hash Normalization (util/hash.rs:40-59)

line_hash(line): trims the line, then collapses any run of ASCII whitespace to a single space while hashing byte-by-byte. This makes anchors immune to:

  • Leading/trailing whitespace
  • Tab vs space differences
  • Multiple consecutive spaces

While still distinguishing actual content differences.

Letter Encoding (util/hash.rs:70-79)

#![allow(unused)]
fn main() {
pub fn encode_hash(hash: u32, len: usize) -> String {
    assert!(len > 0 && len <= 4);
    let mut result = String::with_capacity(len);
    for i in 0..len {
        let byte = ((hash >> (i * 8)) % 26) as u8 + b'a';
        result.push(byte as char);
    }
    result
}
}

Each output letter comes from a different byte of the u32, mod 26, mapped to 'a'..'z'. Default hash_len = 3 → 26³ = 17,576 possible values per anchor component. ParsedAnchor::parse rejects any anchor whose segments aren’t all-lowercase ASCII.


The Three Schemes

A. ContentOnly (scheme.rs:192-276)

Anchor format: LINE:LOCAL (e.g. 22:abc)

Hash computation: encode_hash(line_hash(line), hash_len) for that single line.

Context: None. Validation reads only the anchored line.

Properties:

  • Edits above/below never invalidate (only the line’s own content matters)
  • Cheapest to validate (1 line hashed)
  • Highest collision risk (short lines like }, blank lines, }; all hash identically)
  • Lowest anchor churn after edits
  • validation_window_lines = 1

B. ChunkFingerprint (scheme.rs:278-421)

Anchor format: LINE:LOCAL:CHUNK (e.g. 22:abc:rst)

Hash computation:

  • LOCAL = per-line hash (same as ContentOnly)
  • CHUNK = fold of all line hashes in a fixed-size, page-aligned chunk:
#![allow(unused)]
fn main() {
chunk_start = (line_idx / chunk_size) * chunk_size;
combined = fnv1a_32(b"chunk");
for each line in chunk:
    combined ^= line_hash(line);
    combined = combined.wrapping_mul(16_777_619);
encode_hash(combined, hash_len)
}

Default chunk_size = 8 (production config); all lines in the same chunk share one context fingerprint.

Properties:

  • Any edit anywhere in the chunk invalidates every anchor in that chunk (“collateral staleness”)
  • Reduces false “still valid” acceptance vs ContentOnly (two independent hash checks)
  • validation_window_lines = chunk_size (default 8 lines re-hashed per validation)
  • Explicitly rejects anchors that omit the context — refuses silent degradation to ContentOnly
  • Shift recovery works well for shifts aligned to chunk boundaries (content of destination chunk is identical)

C. CheckpointChain (scheme.rs:423-559)

Anchor format: LINE:LOCAL:CKPT (e.g. 22:abc:rst — same shape as B, different semantics)

Hash computation:

  • LOCAL = per-line hash
  • CKPT = running chain from the nearest checkpoint boundary through the current line:
#![allow(unused)]
fn main() {
checkpoint_start = (line_idx / checkpoint_interval) * checkpoint_interval;
chain = fnv1a_32(b"ckpt");
for each line from checkpoint_start..=line_idx:
    chain ^= line_hash(line);
    chain = chain.wrapping_mul(16_777_619);
encode_hash(chain, hash_len)
}

Default checkpoint_interval = 32.

Properties:

  • Position-sensitive: two identical lines at different offsets from checkpoint boundary get different fingerprints
  • Any edit at or above the line (within the checkpoint window) invalidates the anchor
  • Best collision resistance for repeated content (position distinguishes duplicates)
  • Highest anchor churn — pure line shifts almost always invalidate
  • validation_window_lines grows with distance from checkpoint (average ~16, worst case 32)
  • Not shipped in production — fully implemented and tested but not reachable from build_scheme() in config

Cross-Scheme Comparison

PropertyContentOnlyChunkFingerprintCheckpointChain
Hash cost per validation1 lineup to 8 linesup to 32 lines (avg ~16)
Edits above anchorNever invalidateOnly if in same chunkAny edit in checkpoint window invalidates
Distant unrelated editsImmuneImmune (outside chunk)Immune (outside window)
Token overhead / line:abc (4 chars):abc:rst (8 chars):abc:rst (8 chars)
Collision probabilityHighestLower (two checks)Lowest (position-sensitive)
Shift recovery successBest (local hash often unique)Good if chunk-alignedWorst (chain breaks on any shift)
Anchor churn after editsLowestMediumHighest
Production statusYes ("content_only")Yes ("chunk", default)Benchmark-only

The Read/Edit Round Trip

Read → Anchored Output

read_file.rs:27-68, format_hashline_content:

  1. Splits the full file into lines (anchors need whole-file context for chunk/checkpoint computation)
  2. Calls scheme.generate_anchors(&all_lines)
  3. For the requested offset/limit window, renders each line as:
    {line_num}:{local}:{ctx}→{content}
    
    Using Unicode → as the anchor/content separator.

Grep → Anchored Results

grep.rs: Runs standard ripgrep, then inject_anchors rewrites:

  • Match lines: 123: let x = 1; → 123:abc:rst: let x = 1;
  • Context lines: 124- ... → 124:abc:rst- ...

Per-file anchors are cached in a HashMap<PathBuf, Vec<Anchor>> for the call.

Edit → Anchor Validation and Apply

edit/apply.rs, validate_anchor (lines 531-668):

  1. Strip residual content: removes any trailing →content or ->content the model may have copied from read output
  2. Parse: ParsedAnchor::parse; if that fails, tries recover_anchor_by_suffix (when exactly one line’s hash matches a dropped line number)
  3. Validate: calls scheme.validate(&parsed, lines):
    • Valid → proceed
    • OutOfRange → AnchorNotFound error
    • Stale → calls scheme.find_shifted(...), builds rich error with shifted_to/shifted_anchor/ambiguous_candidates plus a fresh-anchored context snippet

All ops validate against the same pre-edit snapshot. Valid ops are sorted bottom-up (higher line numbers first) and spliced to avoid interference. The tool returns a fresh-anchor snippet around the edited region for immediate follow-up edits.


Shift Recovery (find_shifted_generic, scheme.rs:571-620)

Shared by all three schemes:

  1. Scan ±search_radius lines (default DEFAULT_SEARCH_RADIUS = 15) around original line
  2. Filter candidates by local-hash match
  3. For schemes with context: re-validate full scheme at each candidate
  4. Results: 0 candidates → NotFound, 1 → Found{new_line}, ≥2 → Ambiguous{candidates}

Recovery characteristics per scheme:

  • ContentOnly: best — local hash alone often uniquely identifies the line
  • ChunkFingerprint: good for chunk-aligned shifts (content of destination chunk is identical); poor for arbitrary shifts
  • CheckpointChain: worst — chain breaks on almost any shift; falls back to local-hash-only matching which is ambiguous for repeated content

Configuration Scope

HashlineSchemeParams (config.rs:15-31):

  • Fields: scheme ("chunk" default or "content_only"), hash_len (default 3), chunk_size (default 8)
  • Registered as a ResourceType per tool-server session
  • Per-session, uniform across read/edit/grep — not per-file, not per-tool-call
  • Hard mutual exclusion: a session cannot mix standard file tools with hashline tools — it’s all-or-nothing

Only "content_only" and "chunk" are accepted by build_scheme(). CheckpointChain has no config string and lives only in the benchmark harness.


Architectural Insight: Why This Design?

The hashline system solves a specific problem that exact-match editing cannot: robust addressing in files with repeated patterns. A file with 20 occurrences of return null; cannot be addressed by exact string match without additional context. Hashlines solve this by giving each line a position-aware fingerprint.

The tradeoff is clear:

  • More tokens per read (anchors add 4-8 chars per line)
  • More protocol complexity (model must understand anchor format)
  • Collateral staleness (nearby edits can invalidate unrelated anchors)

But in exchange:

  • No ambiguous match failures (the core failure mode of exact-match editing)
  • Atomic batch semantics (stale anchor → reject all → retry with fresh state)
  • Shift-tolerant addressing (find_shifted can recover without re-reading)
  • Self-healing error messages (errors include fresh anchors for immediate retry)

The production choice of ChunkFingerprint as default represents a middle ground: more robust than ContentOnly (catches stale references from nearby edits) without the extreme churn of CheckpointChain (which invalidates on any upstream change).