← Back to Insights
Briefing · Evidence Governance

Cheap Analysis Creates an Expensive Proof Problem

When agents can generate findings faster than organizations can govern definitions, denominators, methods, and review, error scales with insight.

Date  August 2026 Prepared as  Control test ✓ Verified  16 citations checked
Conditional sign-off verdict

Accelerate exploration, but do not promote consequential findings without proof. Every promoted result should carry its intended population, denominator, time grain, complete run lineage, and an independent reproduction from authoritative data.

The sign-off test

Owner

Who owns the decision, the business definition, the authoritative source, and the review threshold?

Briefing

Which prompts, models, hypotheses, filters, joins, transformations, populations, and denominators produced the claim?

Proof

Can an independent analyst reproduce the promoted result from frozen code or query and authoritative data?

What leaders should take from it

1
A thousand analyses are not a thousand replications.

Autonomous analysts can take the same fixed data and question to conflicting conclusions. Unreported analytic freedom turns one selected result into an evidence-selection problem.

2
The plausible answer may use the wrong population.

Documented failures include annual income averaged into fortnightly welfare debts and a data agent reporting 5,062,338 users where the reported real number was 800 million.

3
A named reviewer is not the same as an effective review.

AI-supported citation errors escaped internal review in a government assurance report. External scrutiny triggered line-by-line checking, escalated assurance, and a partial refund.

4
The controls exist, but their coverage is narrow.

Prespecification, multiplicity adjustment, semantic context, lineage, result-set evaluation, independent challenge, and accountable use-case registers are deployed in bounded regimes. Routine internal findings can still sit outside them.

Where the evidence stops

No public source found in this review measures machine-generated findings, promoted findings, independent reproductions, reversals, near misses, and escaped errors on one enterprise-wide denominator. The honest verdict is mostly tolerated and selectively managed, not a measured epidemic.

The evidence verdict

The risk is mostly tolerated and selectively managed. Strong controls exist in clinical trials, bounded banking-model practice, public-sector AI policy, and mature data-agent deployments. The gap is ordinary self-service analytics: findings can be generated outside those regimes, and public evidence does not show what share is independently reproduced before action.

The operating boundary

Explore cheaply

Let agents test hypotheses, transformations, models, and filters. Keep the output visibly non-decision-grade.

Package evidence

Attach the estimand, population, denominator, time grain, source, code or query, complete run lineage, and known alternatives.

Promote deliberately

Require independent reproduction when a finding becomes a decision, publication, model input, or operating change.

Five findings

01
A thousand analyses create selection risk, not replication.

A peer-reviewed study produced 4,946 autonomous analyses across three fixed tasks. Even among the 3,303 runs that passed an LLM-auditor compliance screen, results varied materially, and model or persona choice shifted the support distribution. The exact false-positive rate for real agent runs is unknown because runs are correlated, but a selected result is uninterpretable when the analytic search is hidden.

02
A technically valid derivation can answer the wrong question.

Robodebt averaged annual income into fortnightly welfare-debt estimates despite the source lacking the necessary time detail. A separate production case study records an AI agent reporting 5,062,338 users where the reported real number was 800 million. Population, grain, filter, and business definition are load-bearing parts of the answer.

03
Review fails when it is a title rather than a reproduction duty.

AI-supported citation errors escaped internal review in a government assurance report, then triggered line-by-line checking, escalated assurance, and a partial refund after external discovery. A prose read cannot test table choice, joins, denominator, or undisclosed analytic freedom.

04
The control set is known, deployed, and bounded.

Institutions use prespecification, multiplicity correction, governed definitions, lineage, result-set evaluations, permission enforcement, independent challenge, accountable officials, registers, and impact assessments. Those controls do not yet form one broad regime for ordinary generative or agentic analytics.

05
Successful execution is a weak correctness proxy.

Peer-reviewed bioinformatics evidence shows correctness falling sharply as task complexity rises, even after self-correction. A later benchmark reported no tested model above 40% against reconstructed human baselines. Both are bounded studies, not universal failure rates, but they reject execution success as proof of a valid estimand or conclusion.

What to do differently

01
Separate exploration from decision-grade analysis.

Declare the promotion boundary and its consequence threshold before scaling self-service access.

02
Log the full analytic search.

Record every model, prompt, hypothesis, transformation, filter, and result. Prespecify the primary test or report the specification curve and apply appropriate multiplicity control.

03
Put a denominator card beside every material number.

Name the population, excluded cases, entity type, units, geography, stock or flow, time grain, and whether the source can support the inference.

04
Review by reproduction.

Give an independent analyst the authoritative data and frozen query or code. Compare the rerun, do not merely read the narrative.

05
Build governance into the product.

Ship semantic contracts, human-confirmed definitions, source lineage, result-set evaluations, linked assumptions, permission enforcement, and an auditable query trail with the agent.

06
Measure control capacity on one denominator.

Track generated, promoted, independently reproduced, reversed, near-miss, and escaped findings, plus review age and analyst hours.

The frontier measure

As machine-generated finding volume rises, what fraction of consequential findings are independently reproduced from authoritative data before action, and does that fraction fall? This one ratio would distinguish managed risk from tolerated risk.

Verification ledger

16
Checked
all cited sources tested against the claims used
0
Fabricated
no invented source survived the verification pass
7
Corrected
claims narrowed or metadata repaired
1
Demoted
preprint evidence kept below the headline
ConfirmedMany AI analysts: 4,946 autonomous runs across three tasks; 3,303 passed an LLM-auditor compliance screen, yet compliant findings still varied and were steerable by model and persona.pnas.org
ConfirmedFalse-Positive Psychology: 60.7% of 15,000 simulated null samples produced at least one p below .05 under four combined flexibilities. This is scenario-specific, not a literature-wide falsehood rate.sagepub.com
CorrectedFDA multiplicity guidance: final nonbinding guidance recommends prospective endpoint and multiplicity control for covered clinical trials.fda.gov
ConfirmedCommonwealth Ombudsman: annual income data lacked the detail needed for fortnightly entitlement, existing guidance warned of error, and no over-calculation modelling had been done.ombudsman.gov.au
ConfirmedRobodebt Royal Commission: A$746 million reimbursed to about 381,000 people and A$1.751 billion in debts written off.royalcommission.gov.au
CorrectedAustralian ministerial account: 416,000 Australians were issued unlawful debts; its rounded financial totals differ from the Royal Commission accounting.dss.gov.au
CorrectedOpenAI data-agent case: the 5,062,338 versus 800 million error came from an unnamed agent. Context, evaluation, lineage, and permission controls are self-reported vendor evidence.open-metadata.org
CorrectedAustralian parliamentary record: internal review missed errors; external discovery led to line-by-line checks, chief-risk-officer and board assurance, and A$97,587.11 recovered.aph.gov.au
ConfirmedDEWR replacement report: the current version discloses Azure OpenAI GPT-4o use in the technical methodology and replaced an earlier version again in February 2026.dewr.gov.au
ConfirmedMata v. Avianca: lawyers submitted nonexistent cases and fake quotations after relying on ChatGPT and failing their professional verification duty.uscourts.gov
CorrectedPLOS ONE bioinformatics study: correctness fractions of 60%, 88%, 25%, 13%, and 0% cover all self-correct outputs across five manually assigned complexity levels.plos.org
DemotedIDA-Bench: no tested model exceeded 40%, but this is a 25-task preprint with simulated users, reconstructed baselines, and dated model snapshots.arxiv.org
CorrectedFederal Reserve model-risk guidance: effective challenge is named as sound practice, but the guidance is nonbinding and expressly excludes generative and agentic AI.federalreserve.gov
ConfirmedNIST AI RMF: the voluntary framework names experimental design, representativeness, human oversight, independent review, test documentation, and monitoring. It does not prove adoption.nist.gov
ConfirmedAustralian government AI policy: covered non-corporate Commonwealth entities must use accountable officials, registers, training, operational governance, and in-scope impact assessment.digital.gov.au
CorrectedLLM survey responses: a 43-model NeurIPS study found order and label biases that undermine the studied population-alignment method, not every simulated-respondent use.neurips.cc
What would change this conclusion

Related work

Control test on verified research · 16/16 citations checked · 0 fabricated · 7 corrected · 1 demoted · 0 unreachable