AI Does Not Modernize a Mainframe. It Can Shorten the Proof.
What measured evidence supports in code translation, documentation recovery, and test generation, and where the risk still concentrates.
It is safe to say yes to an AI-assisted modernization pilot if it is a bounded release with a named business owner, a behavior ledger, independent acceptance tests, reconciliation evidence, and a rollback path. Do not sign off on a portfolio promise built from generated lines of code, successful compilation, or a vendor productivity figure.
The evidence in five moves
GAO's legacy-system evidence points to planning, disposition, cyber risk, interfaces, and delivery control as the durable constraints.
Program-repair studies show plausible patches that later fail manual correctness review. Existing tests are evidence, not semantic proof.
Research supports AI help where outputs can be checked. It does not establish unattended estate conversion or safe cutover.
AI can produce candidate code, explanations, and tests faster than scarce domain experts can validate business behavior unless the evidence pipeline changes too.
Public vendor figures lack the data needed to assess quality, rework, parallel-run cost, or whether the old component actually retires.
The sign-off test
Before funding, ask: can the sponsor name the business flow, data fixtures, interfaces, exceptions, performance bounds, reviewer, acceptance record, rollback, and retirement condition for one release? If not, the next investment is evidence recovery, not mass translation.
Open the Deep Dive: release gates, owner / briefing / proof, and the claim ledger →
There is no public, independent, multi-release study showing that AI reduces total legacy-modernization cost after validation, data conversion, integration, security review, parallel run, rework, and confirmed retirement. IBM's 60% productivity figure is a vendor claim, not independent evidence.