The most important pattern in current AI usage data is one that contradicts a lot of how vendors talk about productivity. The Anthropic Economic Index reports that 52% of Claude conversations augment human work versus 45% that automate it. Augmentation has overtaken automation as the dominant interaction pattern. People are choosing to keep humans in the loop, not because the model cannot run end to end, but because the human plus the model produces better outputs than either alone.
That has direct implications for how a training program should be designed. If the dominant pattern is augmentation, the highest-impact training builds the human’s evaluation and judgment skills, not the human’s prompt-writing speed. Most corporate training programs do the opposite, they teach prompts and skip judgment, then wonder why the productivity numbers do not move.
The leadership pattern in the productivity data
Pair the Anthropic finding with a second one from the McKinsey State of AI 2025. Among organisations getting measurable EBIT impact from AI, high performers are three times more likely than peers to have senior leaders demonstrably owning AI adoption. The leader-engagement effect is large, replicable across industries, and missing in most rollouts. The same report shows only 7% of organisations have scaled AI across the org, and only 39% can attribute any EBIT impact to it. The implementation gap is not a skills problem alone, it is a leadership problem combined with a training problem.
A useful frame: the productivity gain from AI is the product of three things, the tool, the human’s skill at using the tool, and the organisational structure that gets the human time and incentive to use the tool well. Training touches the second. Leadership engagement touches the third. Vendors sell the first. Most rollouts buy the first and skip the other two, then conclude AI does not work.
A training plan that prioritises augmentation
Below is a five-step plan that takes the data seriously. It assumes leadership engagement is real (a named owner, a budget, a quarterly review). Without that, no training plan produces durable results.
Step 1: Pick the augmentation tasks first. For each role in scope, identify three tasks where the AI is a clear collaborator rather than a replacement. A claims handler writing the first draft of a customer letter. A financial analyst drafting variance commentary. A copywriter producing 10 ad variations to choose from. These are augmentation tasks. They are also where the HBS and BCG GPT-4 study reported the largest gains, 12.2% more tasks completed, 25.1% faster, 40% higher quality, on tasks inside the model’s frontier.
Step 2: Match tools to tasks. ChatGPT, Claude, Microsoft Copilot, and Gemini overlap heavily but are not interchangeable. Copilot integrates into Word and Excel and is the lowest-friction option for document and spreadsheet work. Claude tends to outperform on long-form writing and analysis tasks. Gemini integrates into Google Workspace and handles multimodal inputs cleanly. Pick a default per task. Reduce decision fatigue.
Step 3: Run a three-week pilot to build a prompt library. Take 15 to 20 enthusiastic early adopters. Have them work the tasks from step one with the tools from step two for three weeks, keeping a running document of the prompts that produced usable outputs. The library is the artefact. It becomes the spine of the formal training. The pilot also identifies which tasks the model handles well and which it does not.
Step 4: Train evaluation on the prompt library outputs. This is the step most programs skip. For every prompt in the library, the training includes the typical failure mode of the output. The model fabricates citations sometimes. It sometimes overshoots tone for the audience. It sometimes confidently states a number that is wrong. Teach the failure mode alongside the prompt. A copy-paste self-critique prompt every employee should learn:
Read the response above. Identify every factual claim, every named entity, and every numerical figure. For each, rate your confidence on a scale of 1 to 5 and tell me what evidence supports it. Then identify the three claims most likely to be wrong if any of them are wrong.
That single prompt, run consistently after every substantive AI output, builds more durable evaluation skill than any course module on prompt engineering.
Step 5: Measure adoption against task outcomes, not licences activated. The licence-activation metric is a vanity metric. The metrics that matter are task-level: median time to complete a defined output, quality rating on a defined output, hours saved per role per week as self-reported in a quarterly survey. McKinsey’s framing again, “how much EBIT impact did this produce?” If you cannot answer it after two quarters, the program is not working.
Why augmentation training outperforms automation training
The temptation, when the budget is large and the use case looks predictable, is to swing for full automation. Define the workflow, plug in the model, remove the human from the loop. The data on what happens next is unambiguous. The Anthropic data shows users choose augmentation 52% of the time even when automation is available. The HBS and BCG study shows that on tasks outside the model’s capability frontier, the AI-assisted group performed worse than the control. Automation rollouts that do not include a human review step inherit those failures directly. Augmentation rollouts catch them.
The other reason augmentation training wins: it scales. Automation gains plateau quickly because there is a fixed inventory of fully automatable tasks. Augmentation gains compound because the same employee can apply augmentation to a continuously expanding set of tasks. A claims handler who learns evaluation can apply it to letter drafting, then policy summarisation, then complaint handling, then training material review. Each new application produces another increment of productivity.
What the data says about who pulls ahead
The WEF Future of Jobs Report 2025 ranks AI literacy among the top five fastest-growing skill needs through 2030, and projects 39% of workers’ skills as outdated or transformed by then. The high performers in McKinsey’s 2025 cohort are differentiated not by which model they bought but by the discipline of their rollout, named owners, defined cohorts, measurable outcomes. The Anthropic data tells you the dominant pattern in actual usage is augmentation. The HBS and BCG study tells you why training has to include failure modes.
Combine those four findings and the playbook for the next 12 months is unusually clear. Pick the augmentation tasks, match tools to tasks, build the prompt library, train evaluation alongside generation, and measure outcomes at the task level. The companies that run this play seriously will pull ahead. The ones that buy the licences and skip the training will produce the same flat EBIT story McKinsey is documenting in the 61% of organisations that cannot attribute any financial impact to their AI investment. The pattern is visible. The fix is cheap. The reason it is not happening more widely is that most rollouts have a tool budget and not a training one.