← Back to the Overview
Deep Dive · Legacy Modernization

The Evidence Map for AI-Assisted Modernization

Use this to decide whether an AI modernization proposal can enter a controlled release path, not whether its demonstration is impressive.

Gates before you scale

01 · Name the release boundary.

Choose one business flow, source component, data set, interface set, exception class, performance constraint, and retirement condition.

02 · Recover the behavior ledger first.

Link requirements, source routines, data transformations, interfaces, recovered documentation, generated candidates, tests, reviewer, and sign-off evidence.

03 · Use independent acceptance evidence.

Run source-versus-target comparisons, data reconciliations, contract tests, negative cases, and production-like performance checks. Inherited tests are necessary but insufficient.

04 · Measure end-to-end proof throughput.

Track rework, review load, defect escape, reconciliation pass rate, accepted-release lead time, and parallel-run cost beside code output.

05 · Release and retire deliberately.

Make rollback, ownership, and the old-component retirement decision explicit. A translation that creates a permanent parallel estate has not proved economic value.

Owner, briefing, proof

Owner

The accountable business owner decides which outcomes and exceptions must survive, and signs the acceptance record.

Briefing

A release packet names the behavior boundary, baseline, risk, interfaces, evidence gaps, and decision gate.

Proof

Traceable source-to-test-to-reconciliation-to-approval evidence, plus measured post-release behavior and retirement outcome.

Claim ledger

11/11 checked0 fabricated6 corrected3 demoted
Confirmed and qualified: GAO-25-107795 documents 11 agency-reported critical legacy systems and incomplete modernization plans. It supports the program-control problem, not an AI causal claim.
Confirmed: Java program-repair studies show test-suite-adequate patches can be incorrect. They support the validation risk, not a measured mainframe outcome.
Corrected: ICLR code-translation research reports function-level gains from unit-test-filtered self-training. Its 25.5% figure is error-rate reduction, not relative accuracy improvement.
Qualified: IBM's SANER paper reports improvement on samples, not enterprise-system cutover or lifecycle economics.
Demoted: FreshBrew is a public Java-upgrade benchmark with proxy acceptance gates, not mainframe evidence. IBM's 60% productivity result is a vendor claim without public method or independent audit.

What would change our mind

An independently evaluated, multi-release legacy-modernization study that measures AI contribution through validation, data conversion, interface work, security review, parallel run, rework, production outcome, and confirmed retirement. That would answer whether AI reduces total modernization cost or only moves work into the proof queue.

Where this can lead

A sponsor can begin with a release-level diagnostic: owner, behavior boundary, evidence gaps, and acceptance test. If the diagnosis finds a recurring proof gap, build the ledger and release controls across the modernization portfolio.