Reskilling Is Proven in the Work, Not the Certificate
A verified read of the evidence separating workforce capability from training theater: which numbers measure exposure, which measure transfer, and how to test whether an AI program actually changed daily work.
It is safe to say yes to an AI reskilling investment only if the sponsor can show the transfer system behind it: a named workflow with a quality baseline, coached live practice after the course, supervisor follow-through, and a 90-day quality-adjusted measure of whether trained people still run the new method on real work. Completion rates, license counts, prompt volume, and confidence surveys are the starting gate, not the outcome. If the metric can rise while daily work stays unchanged, it is not a reskilling metric.
The four measurement lanes
Completion, licenses, logins, and prompt counts. Necessary funnel data that proves reach and access, never success.
Observed working habits: iterating, checking outputs, knowing when not to use the tool. A process lens for coaching, not a score.
Independently sampled live cases scored for quality, policy, rework, and escalation against the pre-training baseline.
What the program returns after coaching time, manager review load, and rework are counted, not just hours saved.
The sign-off test
Owner
Who owns the target workflow, its quality baseline, and the coaching cadence that turns a completed course into sustained live practice?
Briefing
Which numbers on the program dashboard measure exposure or activity, which measure transfer, and which claims behind them are still contested?
Proof
Can the team produce sampled live cases at day 90 that meet quality and policy thresholds and equal or beat the pre-training baseline?
What leaders should take from it
Decades of peer-reviewed training-transfer research defines success as applying the learned method to real work and sustaining it over time. Certificates, test scores, and satisfaction can show exposure and near-term learning; they do not establish workplace use or performance. This distinction is older than generative AI, and it is the load-bearing test for every reskilling claim.
Meta-analytic evidence ties supervisor, peer, and organizational support positively to transfer and sustainment. The honest reading is association, not causation: a weekly manager check-in is a strong design hypothesis grounded in general transfer evidence, not a demonstrated AI-specific cause of fluency.
In a peer-reviewed field study of 5,172 support agents, AI raised throughput most for less-experienced workers while the most skilled saw little gain and small quality declines. In a survey of nearly 6,000 executives, 69 percent of firms used AI while 89 percent reported no labor-productivity impact over three years. More access or more use cannot stand in for better work.
The AI Fluency Framework's 24 behaviors give programs a concrete vocabulary for observing how people work with AI. The published index built on it observed 11 conversation-visible behaviors on a single vendor platform and is correlational. Treat it as a process lens for coaching, not a validated enterprise capability measure.
Count the share of trained workers who, 90 days later, complete independently sampled live cases through the approved workflow at equal-or-better quality without excess rework or risk. Report 30, 60, and 90 days separately so decay is visible. This is a metric design to publish and test, not an established industry standard.
Three claims run ahead of the evidence: that a vendor-published fluency index measures enterprise skill (the leading one is vendor-primary, single-platform, correlational, and observes fewer than half of its own behavior list), that a specific supervisor cadence causally produces AI capability (support is associated with transfer; no AI-specific causal trial exists), and that the 90-day quality-adjusted transfer rate is an industry standard (it is a defensible synthesis a program should publish and test). The direction of all three is well supported; the strong versions are not.
The Deep Dive holds the action map: the workflow baseline, the course-as-starting-gate design, the supervisor transfer cadence, the four-lane scorecard, the full claim ledger with contested signals, and refresh triggers.
Open the Deep Dive