A fast wrong answer over federated data is worse than a slow right one, because it is believed. This chapter sequences the build so that being slow beats being wrong at every stage.
You now have the full technical picture: annotations, semirings, a typed IR, a validator, and a harness that enforces qualifiers. The remaining question is: in what order do you build this, and what do you watch for once it’s running?
Six levels, each independently useful, each building on the previous:
| Level | Capability | What it gives you |
|---|---|---|
| L0 | Federated exact | List, filter, project, key joins, delegated enforcement. Boring and correct. |
| L1 | Analytical | Metric registry, grain tracking, additivity enforcement, fan-out detection, time semantics. |
| L2 | Ranked retrieval | Search, reranking, fusion, the completeness lattice, aggregate restrictions over approximate results. |
| L3 | Cross-system identity | Link graph, thresholds, confidence propagation, permissioned edges. |
| L4 | Full permission semantics | Guarantee lattice, negation restrictions, rollup subsumption, plan-level authorization and audit. |
| L5 | Optimization | Cost-based planning, semi-join reduction, caching, pre-aggregate routing. |
L0 is table stakes. If you cannot reliably list, filter, and join exact data from your warehouse, nothing else matters. Most teams skip this because it sounds trivial, then discover that their agent is joining on the wrong column half the time. L0 means: deterministic joins on verified keys, with enforcement delegated to the source (let Postgres handle the RLS, don’t re-implement it).
L1 before L2: if your warehouse is your primary asset and your metric definitions are already contested (“which ARR number is the real one?”), start here. The metric registry pays for itself immediately by preventing the sum-across-time class of errors, which are the most common and the hardest to catch after the fact.
L2 before L1: if the unstructured corpus is where the value is (support conversations, documents, customer communications), start with retrieval. The completeness lattice is simpler to implement than the full metric registry and immediately prevents the “12 accounts” class of overconfident counts.
L1 and L2 are independently orderable. Pick based on where your value and your risk concentrate.
L3 after L1+L2: cross-system identity is only needed when you’re actually combining data from different systems. If you’re still querying one warehouse, skip it.
L4 after L3: the full permission semantics (Non-Truman, rollup subsumption) are the most complex and the most organizationally expensive. Do them last because they require buy-in from security and data governance teams, which takes time you can spend on L1–L3.
L5 last, always. Optimization is about making correct answers fast. A fast wrong answer over federated data is worse than a slow right one, because it is believed.
One wrong number costs more than a month of “I don’t know”
When a CRO gets a wrong number from an AI system and discovers it later (and they will — finance will catch it, or the customer will contradict it), trust in the entire system collapses. It does not recover gradually. It recovers by executive mandate, six months later, after a re-evaluation. The cost of “I don’t know” is a follow-up question. The cost of a wrong answer is a roadmap-level setback. Always choose slowness over wrongness.
Use this decision tree:
Most enterprise deployments hit all four eventually. The question is which one you ship first to start delivering value while the rest is built.
The classical inference-control literature assumed a determined human analyst issuing queries by hand. An agent issues a hundred aggregate queries in a loop without getting bored. The economics of the tracker attack changed completely the moment issuing a thousand carefully-varied aggregate queries became free.
# Agent loop: recover individual account ARR from aggregate access
target_account = "A-2201"
total = agent.query("What is total APAC ARR?")
without = agent.query(
f"What is total APAC ARR excluding account {target_account}?"
)
individual_arr = total - without
# The agent just recovered data it shouldn't have at the individual level
@dataclass
class InferenceDefense:
# Layer 1: Query budgets
max_aggregates_per_entity_per_session: int = 10
max_aggregate_queries_per_session: int = 200
# Layer 2: Minimum group size
min_group_size: int = 5 # k-anonymity for aggregates
# Layer 3: Variation detection
max_similar_queries: int = 3 # flag repeated queries with small variations
def check(self, query, session_history) -> Optional[str]:
# Check group size
if query.estimated_group_size() < self.min_group_size:
return "Query rejected: aggregate group too small (min 5)."
# Check budget
entity_queries = session_history.aggregate_queries_targeting(
query.target_entities()
)
if len(entity_queries) >= self.max_aggregates_per_entity_per_session:
return "Query budget exhausted for this entity."
# Check variation pattern
similar = session_history.find_similar(query, threshold=0.9)
if len(similar) >= self.max_similar_queries:
return (
"Multiple similar aggregate queries detected. "
"This pattern may reveal individual-level data. "
"Please reformulate or request direct access."
)
return None
An agent issues these three queries in sequence: (1) “Total ARR for enterprise accounts in APAC”, (2) “Total ARR for enterprise accounts in APAC with more than 100 seats”, (3) “Total ARR for enterprise accounts in APAC with more than 100 seats excluding Acme Corp.” Only one account in APAC has exactly 100–101 seats. What can be inferred from the difference between queries 2 and 3? Would any of the three defense layers catch this?
Query 3 minus Query 2 gives Acme Corp’s exact ARR (assuming they have >100 seats). Layer 2 (min group size) might catch it if the result of Query 3 has fewer than 5 accounts. Layer 3 (variation detection) should catch it: queries 2 and 3 are near-identical with a small exclusion variation — this is exactly the differencing pattern. Layer 1 (budget) would only catch it if the entity “Acme Corp” has already been targeted many times.
Once the system is running, monitor these signals:
metrics:
rejections_total:
labels: [reason_type]
# reason_type: additivity, grain, guarantee, completeness, answerability
# High rejection rate on "additivity" → users are asking time-series
# questions the registry doesn't handle well. Expand the registry.
# High rejection rate on "guarantee" → users are hitting access
# boundaries frequently. Consider expanding access or improving
# the refusal messages to suggest alternatives.
# High rejection rate on "answerability" → users are asking questions
# the available sources cannot answer. Consider adding new sources.
If every answer comes with qualifiers, the system is honest but unhelpful. Track which qualifiers fire most often and invest in improving the underlying data quality:
If the model is generating raw SQL that bypasses the metric registry, track it. Every raw-SQL query is an unvalidated query — it might be correct, but you have no structural guarantee. The goal is zero fallthrough for production answers; raw SQL is acceptable only for exploration and ad-hoc investigation by data engineers.
def track_fallthrough(plan):
if plan.get("op") == "raw_sql":
metrics.increment("raw_sql_fallthrough", labels={
"user_role": plan["principal_role"],
"question_type": classify_question(plan["original_question"]),
})
# Alert if a non-engineer is getting raw SQL answers
if plan["principal_role"] not in ("data_engineer", "analyst"):
alert("Raw SQL answer served to non-technical user")
The registry starts small and grows. Track what fraction of questions hit defined metrics vs. fall through to raw SQL or are unanswerable:
| Category | Target | Action when below target |
|---|---|---|
| Answered via registry | > 80% | Add missing measures to the registry |
| Answered via raw SQL | < 15% | Identify patterns and formalize them |
| Rejected (unanswerable) | < 5% | Add sources or access patterns |
The registry is a living document. Every question that falls through to raw SQL is a signal that a measure or access pattern is missing. The engineering team should have a weekly review of fallthrough queries and a process for promoting common patterns into the registry.
As a design target, every agent response should pass this test:
The forwarding test
If a user forwards this answer to their VP without additional context, will the VP draw the correct conclusion? If not — if the VP would need to know about the top-k bound, the time-aggregation rule, or the access restriction to interpret the number correctly — then the answer must include that information.
The harness enforces this structurally: qualifiers derived from annotations are mandatory in the response. The test is: could someone who didn’t ask the question still interpret the answer correctly?
Three examples of the pattern applied:
# Bad: bare number
"Total ARR at risk: $7.2M across 12 accounts."
# Good: qualified, bounded, contextualized
"At least $2.4M across at least 12 accounts (from the top 50 matching
conversations this quarter — accounts with lower-ranked churn signals
are not included). ARR is the July period-end value. Entity matching
used verified domains only; 4 support organizations could not be linked."
# Good: honest refusal
"I cannot determine which accounts lack an opportunity, because your
access to the opportunities table is restricted. I can tell you which
accounts have open P0 tickets (12 accounts), but cannot assess their
CRM coverage from your access level."
Four areas where the field is actively developing and this book’s framework will need extension:
The metric registry handles semi-additive measures with a time_rule, but real temporal queries are richer: “Show me the trend,” “Compare this quarter to last,” “When did it start declining?” These require a temporal algebra that this book doesn’t cover. The annotation framework extends naturally (each temporal window has its own completeness and grain), but the planning gets complex.
When multiple agents cooperate to answer a question — one retrieves, one calculates, one narrates — the annotations must flow across agent boundaries. The current framework assumes a single validator; extending to distributed validation is an open problem.
Every rejection is a training signal. If the model consistently proposes plans that violate additivity rules, it can be fine-tuned to avoid those patterns. The validator provides a dense, deterministic reward signal — much better than execution-match for RL-based SQL generation. This connection to the reinforcement learning literature (Reasoning-SQL, GRPO-based approaches) is barely explored.
The query budget defense (Section 8.3) is a coarse approximation. Full differential privacy — calibrated noise with a formal privacy guarantee — is the principled solution, but integrating it with the annotation framework requires tracking a privacy budget as another semiring-like quantity that degrades with each query. The shape is right; the details are open.