← Certain Answers
Part IV · Building It · Chapter 8

Sequencing and Defending

A fast wrong answer over federated data is worse than a slow right one, because it is believed. This chapter sequences the build so that being slow beats being wrong at every stage.

You now have the full technical picture: annotations, semirings, a typed IR, a validator, and a harness that enforces qualifiers. The remaining question is: in what order do you build this, and what do you watch for once it’s running?

8.1 The L0–L5 ladder

Six levels, each independently useful, each building on the previous:

LevelCapabilityWhat it gives you
L0Federated exactList, filter, project, key joins, delegated enforcement. Boring and correct.
L1AnalyticalMetric registry, grain tracking, additivity enforcement, fan-out detection, time semantics.
L2Ranked retrievalSearch, reranking, fusion, the completeness lattice, aggregate restrictions over approximate results.
L3Cross-system identityLink graph, thresholds, confidence propagation, permissioned edges.
L4Full permission semanticsGuarantee lattice, negation restrictions, rollup subsumption, plan-level authorization and audit.
L5OptimizationCost-based planning, semi-join reduction, caching, pre-aggregate routing.

Why this order

L0 is table stakes. If you cannot reliably list, filter, and join exact data from your warehouse, nothing else matters. Most teams skip this because it sounds trivial, then discover that their agent is joining on the wrong column half the time. L0 means: deterministic joins on verified keys, with enforcement delegated to the source (let Postgres handle the RLS, don’t re-implement it).

L1 before L2: if your warehouse is your primary asset and your metric definitions are already contested (“which ARR number is the real one?”), start here. The metric registry pays for itself immediately by preventing the sum-across-time class of errors, which are the most common and the hardest to catch after the fact.

L2 before L1: if the unstructured corpus is where the value is (support conversations, documents, customer communications), start with retrieval. The completeness lattice is simpler to implement than the full metric registry and immediately prevents the “12 accounts” class of overconfident counts.

L1 and L2 are independently orderable. Pick based on where your value and your risk concentrate.

L3 after L1+L2: cross-system identity is only needed when you’re actually combining data from different systems. If you’re still querying one warehouse, skip it.

L4 after L3: the full permission semantics (Non-Truman, rollup subsumption) are the most complex and the most organizationally expensive. Do them last because they require buy-in from security and data governance teams, which takes time you can spend on L1–L3.

L5 last, always. Optimization is about making correct answers fast. A fast wrong answer over federated data is worse than a slow right one, because it is believed.

The trust-destruction asymmetry

One wrong number costs more than a month of “I don’t know”

When a CRO gets a wrong number from an AI system and discovers it later (and they will — finance will catch it, or the customer will contradict it), trust in the entire system collapses. It does not recover gradually. It recovers by executive mandate, six months later, after a re-evaluation. The cost of “I don’t know” is a follow-up question. The cost of a wrong answer is a roadmap-level setback. Always choose slowness over wrongness.

8.2 Decision criteria: analytics-first vs. retrieval-first

Use this decision tree:

  1. Does your primary use case involve aggregation over structured data (revenue, metrics, KPIs)? → L1 first (metric registry).
  2. Does it involve finding and summarizing unstructured evidence (conversations, documents, tickets)? → L2 first (completeness lattice).
  3. Does it involve combining data across systems with different ID spaces? → L3 is your critical path; do L0 + L3 and defer L1/L2.
  4. Is your primary concern data governance (who can see what, audit, compliance)? → L0 + L4; defer L1/L2/L3 until the permission model is solid.

Most enterprise deployments hit all four eventually. The question is which one you ship first to start delivering value while the rest is built.

8.3 Differencing attacks in the agent era

The classical inference-control literature assumed a determined human analyst issuing queries by hand. An agent issues a hundred aggregate queries in a loop without getting bored. The economics of the tracker attack changed completely the moment issuing a thousand carefully-varied aggregate queries became free.

The basic attack

# Agent loop: recover individual account ARR from aggregate access
target_account = "A-2201"

total = agent.query("What is total APAC ARR?")
without = agent.query(
    f"What is total APAC ARR excluding account {target_account}?"
)
individual_arr = total - without
# The agent just recovered data it shouldn't have at the individual level

Defense layers

@dataclass
class InferenceDefense:
    # Layer 1: Query budgets
    max_aggregates_per_entity_per_session: int = 10
    max_aggregate_queries_per_session: int = 200

    # Layer 2: Minimum group size
    min_group_size: int = 5  # k-anonymity for aggregates

    # Layer 3: Variation detection
    max_similar_queries: int = 3  # flag repeated queries with small variations

    def check(self, query, session_history) -> Optional[str]:
        # Check group size
        if query.estimated_group_size() < self.min_group_size:
            return "Query rejected: aggregate group too small (min 5)."

        # Check budget
        entity_queries = session_history.aggregate_queries_targeting(
            query.target_entities()
        )
        if len(entity_queries) >= self.max_aggregates_per_entity_per_session:
            return "Query budget exhausted for this entity."

        # Check variation pattern
        similar = session_history.find_similar(query, threshold=0.9)
        if len(similar) >= self.max_similar_queries:
            return (
                "Multiple similar aggregate queries detected. "
                "This pattern may reveal individual-level data. "
                "Please reformulate or request direct access."
            )

        return None
Drill 8.1

An agent issues these three queries in sequence: (1) “Total ARR for enterprise accounts in APAC”, (2) “Total ARR for enterprise accounts in APAC with more than 100 seats”, (3) “Total ARR for enterprise accounts in APAC with more than 100 seats excluding Acme Corp.” Only one account in APAC has exactly 100–101 seats. What can be inferred from the difference between queries 2 and 3? Would any of the three defense layers catch this?

Show answer

Query 3 minus Query 2 gives Acme Corp’s exact ARR (assuming they have >100 seats). Layer 2 (min group size) might catch it if the result of Query 3 has fewer than 5 accounts. Layer 3 (variation detection) should catch it: queries 2 and 3 are near-identical with a small exclusion variation — this is exactly the differencing pattern. Layer 1 (budget) would only catch it if the entity “Acme Corp” has already been targeted many times.

8.4 Operating the system

Once the system is running, monitor these signals:

Rejection rate by annotation type

metrics:
  rejections_total:
    labels: [reason_type]
    # reason_type: additivity, grain, guarantee, completeness, answerability

  # High rejection rate on "additivity" → users are asking time-series
  # questions the registry doesn't handle well. Expand the registry.

  # High rejection rate on "guarantee" → users are hitting access
  # boundaries frequently. Consider expanding access or improving
  # the refusal messages to suggest alternatives.

  # High rejection rate on "answerability" → users are asking questions
  # the available sources cannot answer. Consider adding new sources.

Qualifier frequency

If every answer comes with qualifiers, the system is honest but unhelpful. Track which qualifiers fire most often and invest in improving the underlying data quality:

Raw-SQL fallthrough rate

If the model is generating raw SQL that bypasses the metric registry, track it. Every raw-SQL query is an unvalidated query — it might be correct, but you have no structural guarantee. The goal is zero fallthrough for production answers; raw SQL is acceptable only for exploration and ad-hoc investigation by data engineers.

def track_fallthrough(plan):
    if plan.get("op") == "raw_sql":
        metrics.increment("raw_sql_fallthrough", labels={
            "user_role": plan["principal_role"],
            "question_type": classify_question(plan["original_question"]),
        })
        # Alert if a non-engineer is getting raw SQL answers
        if plan["principal_role"] not in ("data_engineer", "analyst"):
            alert("Raw SQL answer served to non-technical user")

8.5 Metric registry coverage

The registry starts small and grows. Track what fraction of questions hit defined metrics vs. fall through to raw SQL or are unanswerable:

CategoryTargetAction when below target
Answered via registry> 80%Add missing measures to the registry
Answered via raw SQL< 15%Identify patterns and formalize them
Rejected (unanswerable)< 5%Add sources or access patterns

The registry is a living document. Every question that falls through to raw SQL is a signal that a measure or access pattern is missing. The engineering team should have a weekly review of fallthrough queries and a process for promoting common patterns into the registry.

8.6 The honest-answer pattern

As a design target, every agent response should pass this test:

Design principle

The forwarding test

If a user forwards this answer to their VP without additional context, will the VP draw the correct conclusion? If not — if the VP would need to know about the top-k bound, the time-aggregation rule, or the access restriction to interpret the number correctly — then the answer must include that information.

The harness enforces this structurally: qualifiers derived from annotations are mandatory in the response. The test is: could someone who didn’t ask the question still interpret the answer correctly?

Three examples of the pattern applied:

# Bad: bare number
"Total ARR at risk: $7.2M across 12 accounts."

# Good: qualified, bounded, contextualized
"At least $2.4M across at least 12 accounts (from the top 50 matching
conversations this quarter — accounts with lower-ranked churn signals
are not included). ARR is the July period-end value. Entity matching
used verified domains only; 4 support organizations could not be linked."

# Good: honest refusal
"I cannot determine which accounts lack an opportunity, because your
access to the opportunities table is restricted. I can tell you which
accounts have open P0 tickets (12 accounts), but cannot assess their
CRM coverage from your access level."

8.7 What comes after this book

Four areas where the field is actively developing and this book’s framework will need extension:

Temporal semantics

The metric registry handles semi-additive measures with a time_rule, but real temporal queries are richer: “Show me the trend,” “Compare this quarter to last,” “When did it start declining?” These require a temporal algebra that this book doesn’t cover. The annotation framework extends naturally (each temporal window has its own completeness and grain), but the planning gets complex.

Multi-agent composition

When multiple agents cooperate to answer a question — one retrieves, one calculates, one narrates — the annotations must flow across agent boundaries. The current framework assumes a single validator; extending to distributed validation is an open problem.

Learning from rejections

Every rejection is a training signal. If the model consistently proposes plans that violate additivity rules, it can be fine-tuned to avoid those patterns. The validator provides a dense, deterministic reward signal — much better than execution-match for RL-based SQL generation. This connection to the reinforcement learning literature (Reasoning-SQL, GRPO-based approaches) is barely explored.

Differential privacy integration

The query budget defense (Section 8.3) is a coarse approximation. Full differential privacy — calibrated noise with a formal privacy guarantee — is the principled solution, but integrating it with the annotation framework requires tracking a privacy budget as another semiring-like quantity that degrades with each query. The shape is right; the details are open.


Exercises

Implement
  1. Build the monitoring dashboard for your query system. Track: (a) rejection rate by type over time, (b) qualifier frequency, (c) raw-SQL fallthrough rate, (d) query latency by level. Implement alerting rules: alert on >20% rejection rate, alert on raw-SQL answers to non-technical users, alert on queries targeting the same entity more than 10 times in a session.
  2. Implement the differencing-attack detector. Given a session history of aggregate queries, detect the pattern: query A and query B differ by a single exclusion predicate. Flag it, log it, and enforce the query budget. Test with a simulated attack sequence.
Extend
  1. Design the L0→L1 migration for your organization’s data. Identify: (a) the first 5 measures to formalize in the registry, (b) the first grain-mismatch join to protect against, (c) the political obstacle (who currently “owns” these definitions?), and (d) how the technical design routes around the political obstacle. Write the migration plan as a proposal you could present to a data platform team.
  2. The book argues that the same database-theory solutions apply regardless of scale. But at what organizational scale does this system become unnecessary (too few users to matter) or insufficient (too many queries for the validator to handle synchronously)? Characterize both boundaries. At the lower bound, what’s a lighter-weight alternative? At the upper bound, what architectural changes are needed (async validation, probabilistic checking, sampling-based enforcement)?

Further reading