What the productivity studies actually measured
Start with the strongest positive result, because it is stronger than most of what gets claimed and weaker than what gets remembered.
Noy and Zhang (2023) randomised 444 college-educated professionals across two occupation-specific writing tasks: press releases, short reports, analysis plans, delicate emails. Treatment used ChatGPT, control used a placebo tool. Outputs were graded one to seven by blinded professionals in the same occupation.
Time fell from 27 minutes to 17, an effect of −0.83 SD, 95 per cent CI [−1.03, −0.63], p < 0.001. Quality rose from 3.79 to 4.54 out of seven, +0.45 SD, 95 per cent CI [0.27, 0.63], p < 0.001.
That is real, large, and it is the ceiling rather than the norm.
Now the part almost nobody quotes. The authors instrumented the process. Time spent rough-drafting fell from about 50 per cent of the task to 22 per cent, and time spent editing rose from 25 per cent to 53 per cent. Which sounds like people were editing more.
They were not. Sixty-eight per cent of treated participants submitted the model's output without editing it. Average active working time after pasting: three minutes. The correlation between post-paste activity time and final grade was approximately r = 0.00. Treatment grades were roughly equal to the grades raw model output received when graded independently.
The authors' own conclusion: ChatGPT "is increasing productivity primarily by substituting for worker effort."
Not augmenting it. Substituting for it. The human added no detectable value on top of the draft.
Three more results, and they establish who gains.
Customer support, N = 5,172 agents, about 3 million chats (Brynjolfsson, Li & Raymond, QJE, 2025). Resolutions per hour rose 15.2 per cent, from a baseline of 2.18, SE 0.033. Average handle time fell 8.5 per cent. Customer sentiment rose 17.7 points. Net promoter score did not move.
The heterogeneity is the finding. Lowest-skill quintile: about +36 per cent. Highest-skill quintile: approximately zero, with small declines on resolution rate and satisfaction. Agents with under one month of tenure gained most; agents over twelve months gained nothing. The authors' proposed mechanism is that the tool disseminates the tacit knowledge of the firm's best agents to everyone else.
Read what that implies. The value came from codifying what high performers already did. The high performers got nothing. And the authors flag, explicitly, that they cannot measure what happens to the training signal if the best workers begin relying on the system that was built from their behaviour.
Software development. A laboratory trial of 70 developers building an HTTP server found completion time falling from 160.89 minutes to 71.17, a 55.8 per cent reduction. The confidence interval is quoted almost nowhere: [21 per cent, 89 per cent] (Peng et al., 2023). Across three field experiments with 4,867 developers at real firms, pull requests completed rose 26.08 per cent, SE 10.3 per cent, with no individual firm reaching significance on its own (Cui et al., Management Science, 2025).
And the population where the sign flips. Sixteen experienced open-source developers, 246 real tasks they had nominated as valuable, in repositories averaging over a million lines that they had worked in for five years, randomised task by task. AI access increased completion time by 19 per cent, 95 per cent CI approximately [2 per cent, 36 per cent] (METR, 2025).
Screen recordings show where it went: less time coding, less time reading and searching, more time prompting, about 4 per cent of total time waiting on generations and about 9 per cent reviewing and cleaning output. Developers accepted fewer than 44 per cent of generations. All of them reported needing to modify what came back.
Line the four up and the pattern is not "AI helps" or "AI does not help". Every study where the population was less expert than the task found gains. The one study where the population was more expert than the task found a loss. Before asking whether AI will help with a task, ask which side of that line you sit on for this task, because the sign of the effect follows it.
That is the shape of the next lesson.
What does the Noy and Zhang result actually establish?
Sixteen experienced developers were 19 per cent slower with AI on their own repositories. What is the most defensible reading?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.