Proxy Metrics Are Theater Unless Finance Can Trace the Value
A client-ready view of what survives AI ROI scrutiny: ledger lines, task fit, the gain-to-P&L chain of custody, and the difference between efficiency bets and positioning bets.
It is safe to claim AI ROI only when the metric traces to finance: cost displaced, headcount not backfilled, or unit-cost cycle time on a process with a known baseline. Hours saved, adoption rate, perceived productivity, and vendor ROI multiples stay internal signals until finance can reconcile them to the general ledger.
The measurement ladder
Usage, prompts, licenses, demos, and adoption. Useful for change tracking, not ROI.
Speed, quality, or volume improves on bounded work. Real evidence, but still upstream of finance.
The gain survives review, integration, governance, and process change inside the operating workflow.
Cost, headcount, unit cost, margin, or revenue shows up against a finance-owned baseline.
Owner, briefing, proof
Owner
Finance-owned baseline and value-recognition rule, not an AI-team estimate.
Briefing
Efficiency or positioning decision brief, with the measurement standard named before the pilot starts.
Proof
A ledger bridge that separates task gain, workflow change, cashed savings, redeployed capacity, and shadow costs.
What leaders should take from it
Finance can audit hard cost removed, headcount not backfilled, or unit-cost cycle time. Proxy metrics become theater when they never touch an income statement.
Controlled studies show strong gains on bounded work and negative effects on some expert work. The task frontier matters more than the headline number.
Published evidence supports task gains, but not a peer-reviewed, finance-controlled enterprise ROI multiple. Precise vendor ROI figures should not carry the argument.
Attribution, redeployed hours, hidden governance costs, integration work, and review bottlenecks break the chain between a pilot win and booked value.
Infrastructure vendors, consultancies, and acquired observability tools often get paid before the client proves ROI. Treat their numbers as claims to test, not proof.
Three claims run ahead of the evidence: that the MIT NANDA 95% figure is an audited failure rate, that any single productivity effect generalizes across tasks, or that a vendor ROI multiple proves enterprise value. The defensible claim is narrower and more useful: AI ROI survives only when finance can trace it.
First moves before hiring anyone
Ask the controller to define the before state, counterfactual, denominator, and ledger line. If the AI team owns the baseline, mark the result as an internal signal.
Efficiency plays need hard-dollar discipline. Positioning bets need option-value framing and explicit exemption from quarterly ROI promises.
When hours saved are claimed, separate cost actually removed from time reinvested into more work. Most productivity claims die in that second column.
Governance, system integration, change management, token spend, security review, and support all belong in the business case before approval.
Compare the client against its own full cohort of attempts, not industry averages that exclude failed or abandoned deployments.
Start with one AI initiative, one finance baseline, and one proof path. If the gap is material, widen to a readiness look at AI value measurement, and build the operating cadence only when the sponsor wants it run.
Claim ledger
- A peer-reviewed study causally identifies firm-level AI ROI.
- The DellAcqua/METR heterogeneity pattern is contradicted or fails to replicate.
- A primary source corrects, replicates, or retracts the MIT NANDA 95% figure.
- An independent audited enterprise-ROI standard gains adoption.
- A large enterprise publishes finance-controlled, GL-reconciled AI ROI with a real holdout.