Structured Policy Analysis
Who Captures AI's Productivity Gains
Who captures AI's productivity gains, novices or experts, and is in-task compression durable upskilling or fragile dependence. AI research grounded in evidence, structured by causal mechanisms. Independent verification required.
Key Findings
Research suggests generative AI often raises measured performance most for novices and below-average workers on tasks inside its competence. In a call center, the least experienced agents gained around 34 percent while the most experienced gained little. The likely mechanism is that the tool codifies and broadcasts the tacit patterns of top performers, so AI can substitute for experience that newer workers have not yet built. This pattern is not universal. On tasks outside the model's reliable frontier, AI can make workers more likely to be wrong, and in some field settings the strongest performers gained while the weakest were hurt. Whether in-task compression becomes durable upskilling or fragile dependence appears to depend on deployment design, and some studies find skills erode when AI is used as a crutch.
Most evidence is recent, short-horizon, and setting-specific. In-task performance compression is not the same as durable skill gain or wage compression. Findings from one task, tool generation, or worker population do not necessarily generalize.
Novices gain most where AI is reliable
Multiple field experiments find the largest measured gains for less experienced or lower-skilled workers in customer support, coding, taxi dispatch, and resume writing. The common thread is that AI encodes practices the novice has not yet learned.
The pattern can invert outside the frontier
On tasks beyond the model's reliable range, consultants using AI were about 19 percentage points less likely to be correct. In a Kenyan entrepreneur study, high performers gained while low performers did worse, the opposite of compression.
Compression is trust-fragile
Gains depend on workers correctly judging when to trust AI. Studies of clinicians show incorrect AI advice can drag down accuracy regardless of expertise, and over-reliance can overturn initially correct judgments.
Durability depends on how AI is deployed
High school students who used unguarded AI as a crutch scored worse without it later. A tutor-mode design with guardrails largely removed that harm. Experienced endoscopists detected fewer adenomas without AI after routine AI exposure.
Some gains do persist for lower-ability workers
An open-source software study found lower-ability maintainers benefited more and the effect persisted for about two years, with task reallocation toward coding. Persistence varies by setting and is not guaranteed.
In-task compression is not wage compression
Who keeps the surplus is a separate question. Industry data shows large wage premiums for advanced AI skills and concern that gains may accrue to capital, which can run opposite to the within-task leveling effect.
Research Findings
Sources
What this means in practice
Work where AI compresses the novice-expert gap often involves manually pulling together context from many records, drafting responses or analyses by hand, and tracking outcomes to see what good performance looks like. These processes are typically handled with systems that automate the repetitive parts so output is consistent regardless of who does the task.
- Ingest records, prior cases, and reference material into one place
- Automate drafting, modeling, and tracking so routine steps are consistent
- Generate clear, repeatable outputs and reports for review
Related Research
Does AI Coding Assistance Actually Speed Developers Up?
When AI coding assistance speeds developers up versus slows them down, and why developers are often wrong about which is happening
AI Tutoring Outcomes
Do AI tutoring systems improve student learning compared with business-as-usual instruction and high-dosage human tutoring?
RAG and Hallucination Insights
Why retrieval-augmented generation reduces but does not fix LLM hallucination, split by retrieval failure vs faithfulness failure
Do Structured Interviews Actually Predict Job Performance?
After the 2022 validity reassessment, how much does a hiring decision rest on an interview, and does structure work by adding insight or by reducing interviewers' own judgment error?