There are two numbers from McKinsey’s State of AI 2025 that, taken together, explain almost everything wrong with the typical corporate AI rollout. 88% of organisations report using AI in at least one function, up 10 points year on year. Only 7% say AI is fully scaled across the org, and only 39% can attribute any EBIT impact to it. Most of those 39% report under 5% EBIT contribution. The other 61% have spent the budget without producing the result.
This is the implementation gap. It is the single most important context for thinking about AI training, and it is missing from almost every training pitch deck. The mistake is not under-investing in tools. The mistake is under-investing in the human integration work that turns the tool into a measurable outcome.
What the data actually shows works
Strip the marketing layer off the rollout literature and a small set of practices keep recurring in the high-performer cohort. McKinsey’s 2025 high-performer analysis frames it bluntly, the organisations getting EBIT impact from AI are three times more likely to have senior leaders demonstrably owning adoption. The other consistent differentiator is that high performers are systematic about training, defined cohorts, named owners, measurable outcomes.
Three findings worth pinning down before you design a program:
The HBS and BCG GPT-4 study found 758 consultants completed 12.2% more tasks and 25.1% faster on tasks inside the model’s frontier, but performed worse than the control on out-of-frontier tasks. Training that includes “what AI does badly on this task” outperforms training that only shows “what AI does well.”
The Anthropic Economic Index shows augmentation (52%) has overtaken automation (45%) as the dominant interaction pattern. Workers are choosing to keep humans in the loop. Training that assumes “AI does the task end to end” does not match how people actually use the tools.
The WEF Future of Jobs Report 2025 ranks AI literacy among the top five fastest-growing skill needs through 2030. Demand for the skill is outrunning supply.
A training program that closes the implementation gap
Below is a five-step structure for an internal program that produces measurable adoption rather than completed-course-counts. None of these steps are exotic. Most rollouts skip three of them.
Step 1: Pick the work, not the tool. Identify three to five specific tasks per role where AI could realistically help. A claims handler summarising a long incident report. A financial analyst building a variance commentary. A copywriter producing 10 ad variations. These tasks become the training content, not “an introduction to ChatGPT.”
Step 2: Match the tool to the task. ChatGPT, Claude, Microsoft Copilot, and Gemini overlap heavily, but each has tasks where it tends to outperform. Copilot inside Word and Excel reduces friction for document and spreadsheet workflows. Claude tends to outperform on long-form writing and analysis tasks above 50,000 tokens. Gemini integrates into Google Workspace and handles multi-modal inputs cleanly. Pick deliberately rather than by default.
Step 3: Build a prompt library before you build a course. Run a three-week pilot with the most enthusiastic 15 to 20 employees in your organisation. Collect every prompt that produced a good output. By the end of the pilot you have an internal prompt library, organised by task, that becomes the spine of the formal training. This is what new joiners learn from. It is also what the laggards copy and paste when they want to try the tool but do not know where to start.
Step 4: Teach evaluation, not just generation. Most AI training is 80% prompt engineering. The higher-impact 80% is teaching people how to spot a confidently wrong AI output. The single most useful exercise is a 30-minute weekly session where each team member brings one AI output and the group identifies what could be wrong with it. That session builds more durable skill than a one-off two-day workshop.
Step 5: Measure adoption against task outcomes, not licences activated. A licence activated is a vanity metric. The right metrics are task-level: median time to draft a customer email before and after, error rate on a defined output, hours saved per role per week as self-reported in a quarterly survey. McKinsey’s framing is the right one here, “how much EBIT impact did this produce?” If you cannot answer it after two quarters, the program is not working.
A copy-paste prompt for the evaluation habit
The hardest skill to embed in a training program is the discipline of evaluating an AI output before sending it. A repeatable prompt pattern helps:
Read the response below. List every factual claim, every named entity, and every numerical figure. For each, rate your confidence on 1 to 5 and explain what evidence supports it. Then identify the three claims most likely to be wrong if any are wrong, and tell me how to verify each one in under five minutes.
Use this on every substantive AI output before it ships externally. It does not eliminate errors. It surfaces the highest-risk ones so the human knows where to look. Combined with a quick verification step in Perplexity for any claim above the team’s risk threshold, the error rate drops sharply.
The honest conclusion
Most corporate AI rollouts are running at the 88% adoption level and the 7% scaled level simultaneously. The training fix is not more workshops. It is a tighter loop between the work, the tool, the prompt library, and the evaluation habit, with named owners and measurable outcomes. The companies that close the implementation gap in the next 12 months will not be the ones that bought the most licences. They will be the ones that built the smallest, most disciplined training program around the tasks that actually matter, then expanded once the metrics moved. That is the unglamorous version of “AI transformation.” It is also the version that survives the next budget cycle.