The Evidence Map for Reskilling That Holds
The action layer behind the core verdict: how to run an AI reskilling program as a workflow-transfer system, and how to measure it without overstating what the evidence supports.
Use this before funding, renewing, or reporting on an AI training program. The course content is rarely the problem; the missing piece is the system around it: a named workflow, coached live practice, quality scored separately from activity, and a sustainment check months later. Treat completion as the starting gate, make the workgroup's reviewed method of work the unit of reskilling, and design the measurement so it earns trust instead of reading as surveillance.
The full findings
Peer-reviewed training-transfer meta-analyses define success as generalization of learned behavior to the job plus maintenance over time. Certificates, satisfaction, and test scores show exposure or near-term learning; they do not establish workplace use or performance. No lens in the research disputed this; the argument is over how to create transfer, not whether completion proves it.
Meta-analytic sustainment correlations run .51 for supervisor support, .48 for peer support, and .32 for organizational support, with overlapping intervals and no reliable winner among support types. A second meta-analysis confirms supervisor support relates to transfer while showing that same-source measurement inflates the estimates. The safe conclusion: support, opportunity, and reinforcement matter. The evidence does not prove a weekly supervisor check-in causes AI fluency.
In a peer-reviewed study of 5,172 support agents, AI assistance raised resolved chats per hour 15 percent overall and about 30 percent for less-skilled workers, while the highest-skill and highest-tenure groups saw little gain and small declines in resolution rate and satisfaction. A survey of nearly 6,000 executives found 69 percent of firms using AI and 89 percent reporting no labor-productivity impact over three years. These are heterogeneity results, not a universal score, but they corroborate the direction: more access or use cannot stand in for better work.
The AI Fluency Framework supplies four competencies operationalized as 24 behaviors. The published index baseline examined 9,830 substantive conversations and only the 11 conversation-visible behaviors, on one vendor platform, correlationally. Independent peer-reviewed work partly corroborates objective skill measurement: a validated literacy assessment predicted AI-supported task performance where self-rated proficiency did not. That supports measuring skill objectively; it does not validate this specific index.
Numerator: eligible trained workers who, 90 days on, pass at least two independently sampled live cases through the approved workflow at preset quality and policy thresholds, equal-or-better than baseline, without excess rework or risk exceptions. Denominator: trained workers with a fair opportunity to perform. Keep exposure, observed behavior, workflow outcome, and net economics in separate lanes. This is a synthesis consistent with NIST measurement principles, not a validated standard: publish the rubric and test its predictive value.
First moves before hiring anyone
Before training begins, capture the role, the recurring task, the approved AI contribution, the human decision owner, and current quality, cycle time, rework, escalation, and risk rates. Without the baseline, every later number is unanchored.
Require a safe practice case and a knowledge check, then move immediately into live work. Completion stays a leading indicator and access prerequisite; it never becomes the success metric.
Weekly for the first month, then biweekly: select one case, inspect the output and its evidence, discuss the human edit, remove one blocker, set the next attempt. Track manager review minutes per learner so the coaching cost stays visible instead of disappearing into the ROI story.
Track eligible opportunities, use, and abandonment for diagnosis. Separately score sampled work for accuracy, completeness, evidence, policy, decision quality, rework, and escalation. Never convert prompt counts into a fluency or performance rank, and score sampled work products with notice and aggregation before any individual scoring: measurement that reads as surveillance costs candor and participation.
Report 30-day transfer, 60-day sustainment, and 90-day sustainment separately so decay is visible, and keep opportunity-to-perform in the denominator so the metric cannot be gamed by shrinking eligibility.
Hold tool, task, and access constant while comparing classroom-only enablement against workflow practice with manager reinforcement. This is the cleanest enterprise test of whether the program, rather than the tool or the participant mix, produced the result.
Owner, briefing, proof
Owner
A named workflow owner plus the supervisors who run the transfer cadence. The program office supplies content and telemetry; the workgroup owns the reviewed method of work, which is the durable unit of reskilling.
Briefing
A one-page read per cohort: the completion funnel, the gap to the first independently passing live case, transfer and decay at 30, 60, and 90 days, manager review load, and which claims stay off the scorecard as contested.
Proof
Independently sampled live cases scored for quality, policy, rework, and escalation against the pre-training baseline, kept as the record that justifies scaling the program or shutting it down.
Claim ledger
- An independent multi-tool, multi-role study validates the 24 fluency behaviors against job performance, risk, or quality outcomes.
- A randomized or strong quasi-experimental enterprise study compares classroom-only AI training with supervisor-led workflow practice at 30, 60, and 90 days.
- The index publisher releases a longitudinal cohort analysis, resolves the sample-year inconsistency, or measures the 13 off-platform behaviors.
- A replicated peer-reviewed study establishes which AI training design improves durable workplace transfer rather than short-term adoption.
- A regulator, standards body, union, or works council sets requirements that materially change employee-level AI-use monitoring or competence assessment.
- A large enterprise publishes audited quality-adjusted reskilling outcomes, including review cost, rework, exceptions, and a comparison baseline.
Start with one workflow and one cohort: write the baseline, run the course as a gate, put the supervisor cadence and sampled quality scoring on that single workflow, and read the 30-60-90 transfer curve before spending further. If the gap between completion and the first independently passing live case is material, widen to a readiness look across roles and workflows. Build the full four-lane scorecard and comparison cohorts only when the sponsor wants the program run as an operating system.