Cheap Analysis Creates an Expensive Proof Problem
When agents can generate findings faster than organizations can govern definitions, denominators, methods, and review, error scales with insight.
Accelerate exploration, but do not promote consequential findings without proof. Every promoted result should carry its intended population, denominator, time grain, complete run lineage, and an independent reproduction from authoritative data.
The sign-off test
Owner
Who owns the decision, the business definition, the authoritative source, and the review threshold?
Briefing
Which prompts, models, hypotheses, filters, joins, transformations, populations, and denominators produced the claim?
Proof
Can an independent analyst reproduce the promoted result from frozen code or query and authoritative data?
What leaders should take from it
Autonomous analysts can take the same fixed data and question to conflicting conclusions. Unreported analytic freedom turns one selected result into an evidence-selection problem.
Documented failures include annual income averaged into fortnightly welfare debts and a data agent reporting 5,062,338 users where the reported real number was 800 million.
AI-supported citation errors escaped internal review in a government assurance report. External scrutiny triggered line-by-line checking, escalated assurance, and a partial refund.
Prespecification, multiplicity adjustment, semantic context, lineage, result-set evaluation, independent challenge, and accountable use-case registers are deployed in bounded regimes. Routine internal findings can still sit outside them.
No public source found in this review measures machine-generated findings, promoted findings, independent reproductions, reversals, near misses, and escaped errors on one enterprise-wide denominator. The honest verdict is mostly tolerated and selectively managed, not a measured epidemic.
The risk is mostly tolerated and selectively managed. Strong controls exist in clinical trials, bounded banking-model practice, public-sector AI policy, and mature data-agent deployments. The gap is ordinary self-service analytics: findings can be generated outside those regimes, and public evidence does not show what share is independently reproduced before action.
The operating boundary
Explore cheaply
Let agents test hypotheses, transformations, models, and filters. Keep the output visibly non-decision-grade.
Package evidence
Attach the estimand, population, denominator, time grain, source, code or query, complete run lineage, and known alternatives.
Promote deliberately
Require independent reproduction when a finding becomes a decision, publication, model input, or operating change.
Five findings
A peer-reviewed study produced 4,946 autonomous analyses across three fixed tasks. Even among the 3,303 runs that passed an LLM-auditor compliance screen, results varied materially, and model or persona choice shifted the support distribution. The exact false-positive rate for real agent runs is unknown because runs are correlated, but a selected result is uninterpretable when the analytic search is hidden.
Robodebt averaged annual income into fortnightly welfare-debt estimates despite the source lacking the necessary time detail. A separate production case study records an AI agent reporting 5,062,338 users where the reported real number was 800 million. Population, grain, filter, and business definition are load-bearing parts of the answer.
AI-supported citation errors escaped internal review in a government assurance report, then triggered line-by-line checking, escalated assurance, and a partial refund after external discovery. A prose read cannot test table choice, joins, denominator, or undisclosed analytic freedom.
Institutions use prespecification, multiplicity correction, governed definitions, lineage, result-set evaluations, permission enforcement, independent challenge, accountable officials, registers, and impact assessments. Those controls do not yet form one broad regime for ordinary generative or agentic analytics.
Peer-reviewed bioinformatics evidence shows correctness falling sharply as task complexity rises, even after self-correction. A later benchmark reported no tested model above 40% against reconstructed human baselines. Both are bounded studies, not universal failure rates, but they reject execution success as proof of a valid estimand or conclusion.
What to do differently
Declare the promotion boundary and its consequence threshold before scaling self-service access.
Record every model, prompt, hypothesis, transformation, filter, and result. Prespecify the primary test or report the specification curve and apply appropriate multiplicity control.
Name the population, excluded cases, entity type, units, geography, stock or flow, time grain, and whether the source can support the inference.
Give an independent analyst the authoritative data and frozen query or code. Compare the rerun, do not merely read the narrative.
Ship semantic contracts, human-confirmed definitions, source lineage, result-set evaluations, linked assumptions, permission enforcement, and an auditable query trail with the agent.
Track generated, promoted, independently reproduced, reversed, near-miss, and escaped findings, plus review age and analyst hours.
As machine-generated finding volume rises, what fraction of consequential findings are independently reproduced from authoritative data before action, and does that fraction fall? This one ratio would distinguish managed risk from tolerated risk.
Verification ledger
- A multi-enterprise audit reports generated, promoted, independently reproduced, reversed, near-miss, and escaped machine findings on one denominator.
- The agentic analytic-multiverse result is independently replicated or overturned with materially different models, datasets, and human-auditor comparisons.
- A public production incident documents a large unprompted analysis set, selective promotion, and a consequential false conclusion.
- A binding control regime extends prespecification, multiplicity, or independent validation to routine generative analytics.
- An independent production audit demonstrates consistently high end-to-end correctness under adversarial population, denominator, grain, and semantic-table perturbations.