← Back to the Overview
Deep Dive · AI Workforce Reskilling

The Evidence Map for Reskilling That Holds

The action layer behind the core verdict: how to run an AI reskilling program as a workflow-transfer system, and how to measure it without overstating what the evidence supports.

Source  Storm Research v2 Verification  15 citations checked Prepared for  Leader discussion
How to use this

Use this before funding, renewing, or reporting on an AI training program. The course content is rarely the problem; the missing piece is the system around it: a named workflow, coached live practice, quality scored separately from activity, and a sustainment check months later. Treat completion as the starting gate, make the workgroup's reviewed method of work the unit of reskilling, and design the measurement so it earns trust instead of reading as surveillance.

The full findings

9/10
Completion is an input; transfer is the outcome.

Peer-reviewed training-transfer meta-analyses define success as generalization of learned behavior to the job plus maintenance over time. Certificates, satisfaction, and test scores show exposure or near-term learning; they do not establish workplace use or performance. No lens in the research disputed this; the argument is over how to create transfer, not whether completion proves it.

8/10
Transfer is manufactured in the work environment.

Meta-analytic sustainment correlations run .51 for supervisor support, .48 for peer support, and .32 for organizational support, with overlapping intervals and no reliable winner among support types. A second meta-analysis confirms supervisor support relates to transfer while showing that same-source measurement inflates the estimates. The safe conclusion: support, opportunity, and reinforcement matter. The evidence does not prove a weekly supervisor check-in causes AI fluency.

8/10
Use quality and use volume are empirically separable.

In a peer-reviewed study of 5,172 support agents, AI assistance raised resolved chats per hour 15 percent overall and about 30 percent for less-skilled workers, while the highest-skill and highest-tenure groups saw little gain and small declines in resolution rate and satisfaction. A survey of nearly 6,000 executives found 69 percent of firms using AI and 89 percent reporting no labor-productivity impact over three years. These are heterogeneity results, not a universal score, but they corroborate the direction: more access or use cannot stand in for better work.

7/10
The 24-behavior fluency taxonomy is a process vocabulary, not an enterprise score.

The AI Fluency Framework supplies four competencies operationalized as 24 behaviors. The published index baseline examined 9,830 substantive conversations and only the 11 conversation-visible behaviors, on one vendor platform, correlationally. Independent peer-reviewed work partly corroborates objective skill measurement: a validated literacy assessment predicted AI-supported task performance where self-rated proficiency did not. That supports measuring skill objectively; it does not validate this specific index.

7/10
The defensible primary metric is sustained, quality-adjusted transfer at day 90.

Numerator: eligible trained workers who, 90 days on, pass at least two independently sampled live cases through the approved workflow at preset quality and policy thresholds, equal-or-better than baseline, without excess rework or risk exceptions. Denominator: trained workers with a fair opportunity to perform. Keep exposure, observed behavior, workflow outcome, and net economics in separate lanes. This is a synthesis consistent with NIST measurement principles, not a validated standard: publish the rubric and test its predictive value.

First moves before hiring anyone

01
Name one workflow and write its baseline.

Before training begins, capture the role, the recurring task, the approved AI contribution, the human decision owner, and current quality, cycle time, rework, escalation, and risk rates. Without the baseline, every later number is unanchored.

02
Treat the course as a starting gate.

Require a safe practice case and a knowledge check, then move immediately into live work. Completion stays a leading indicator and access prerequisite; it never becomes the success metric.

03
Give supervisors a small, explicit transfer cadence.

Weekly for the first month, then biweekly: select one case, inspect the output and its evidence, discuss the human edit, remove one blocker, set the next attempt. Track manager review minutes per learner so the coaching cost stays visible instead of disappearing into the ROI story.

04
Keep volume and quality in different lanes.

Track eligible opportunities, use, and abandonment for diagnosis. Separately score sampled work for accuracy, completeness, evidence, policy, decision quality, rework, and escalation. Never convert prompt counts into a fluency or performance rank, and score sampled work products with notice and aggregation before any individual scoring: measurement that reads as surveillance costs candor and participation.

05
Make the 90-day quality-adjusted transfer rate the primary outcome.

Report 30-day transfer, 60-day sustainment, and 90-day sustainment separately so decay is visible, and keep opportunity-to-perform in the denominator so the metric cannot be gamed by shrinking eligibility.

06
Run a comparison cohort when the stakes justify it.

Hold tool, task, and access constant while comparing classroom-only enablement against workflow practice with manager reinforcement. This is the cleanest enterprise test of whether the program, rather than the tool or the participant mix, produced the result.

Owner, briefing, proof

Owner

A named workflow owner plus the supervisors who run the transfer cadence. The program office supplies content and telemetry; the workgroup owns the reviewed method of work, which is the durable unit of reskilling.

Briefing

A one-page read per cohort: the completion funnel, the gap to the first independently passing live case, transfer and decay at 30, 60, and 90 days, manager review load, and which claims stay off the scorecard as contested.

Proof

Independently sampled live cases scored for quality, policy, rework, and escalation against the pre-training baseline, kept as the record that justifies scaling the program or shutting it down.

Claim ledger

15/15
Checked
citations independently traced to primary sources on July 23, 2026
0
Fabricated
no invented sources or figures surfaced in verification
14
Corrected
scope, attribution, or strength narrowed after source review
5
Demoted
kept as context or contested signal, out of the headline claims
9/10Completion, access, activity, and self-rated confidence are exposure or learning indicators. Durable reskilling requires generalization to live work and maintenance over time; two peer-reviewed meta-analyses carry the claim.Blume 2010 · Hughes 2020
8/10Supervisor, peer, and organizational support are positively associated with transfer and sustainment. The correlations are noncausal, their intervals overlap, and same-source measurement inflates estimates; supervisor cadence remains a design hypothesis, not an AI-specific proven intervention.Hughes 2020 · Blume 2010
8/10Use quality and use volume are separable. AI effects varied sharply by worker skill and tenure in a peer-reviewed field deployment, and firm-level surveys show wide use with little reported productivity impact; the firm data are executive self-reports.QJE 2025 · NBER 34836
7/10The 24 AI Fluency behaviors are a useful process vocabulary, not a validated enterprise outcome score. The published baseline covers 9,830 conversations and 11 conversation-visible behaviors on one platform; an independent validated assessment (GLAT) supports objective skill measurement generally, not this index.Anthropic 2026 · Jin 2025
7/10A 90-day quality-adjusted sustained transfer rate is the defensible primary metric design. Consistent with NIST principles on defined tasks, operator proficiency, thresholds, benchmarks, and monitoring; it is a synthesis to publish and test, not an established standard.NIST AI 100-1 · 600-1
ContestedVendor fluency-index rates. Vendor-primary, correlational, single-platform, with an unresolved sample-year inconsistency; quote exact figures only with attribution and the report's stated limitations, never as an enterprise capability score.Anthropic 2026
ContestedBrief AI training can improve performance. A randomized student study found a 9.5-minute intervention raised adoption and grades; a workplace preprint found an imposed pairing protocol associated with worse output under a confounded design. Durability and workplace transfer are untested in both.arXiv preprints
ContestedManager support predicts frequent AI use. Large cross-sectional survey evidence associates workflow integration and manager support with frequent use; frequency is not work quality, and the design cannot show causation.Gallup 2026
What would change our mind
Where to start

Start with one workflow and one cohort: write the baseline, run the course as a gate, put the supervisor cadence and sampled quality scoring on that single workflow, and read the 30-60-90 transfer curve before spending further. If the gap between completion and the first independently passing live case is material, widen to a readiness look across roles and workflows. Build the full four-lane scorecard and comparison cohorts only when the sponsor wants the program run as an operating system.

Deep Dive staged from verified Storm Research v2 · nothing here asserts above the registry calibration