When two is worse than one
The finding that organises the rest of the course.
Vaccaro, Almaatouq and Malone (Nature Human Behaviour, 2024) systematically reviewed 74 peer-reviewed papers, containing 106 experiments and 370 effect sizes, published between January 2020 and June 2023. Inclusion required all three of human-alone performance, AI-alone performance, and human-AI performance, each with quantitative measures.
Three quantities, which most summaries collapse into one.
Human-AI synergy, the combination against the better of either alone: Hedges' g = −0.23, 95 per cent CI [−0.39, −0.07], p = 0.005. Significantly worse.
Human augmentation, the combination against the human alone: g = 0.64, 95 per cent CI [0.53, 0.74], p < 0.001. Substantially better.
AI augmentation, the combination against the AI alone: g = 0.30, 95 per cent CI [−0.03, 0.62], p = 0.072. Not significant.
In one sentence: giving a person an AI reliably makes that person better than they were, and does not reliably produce something better than the best available single agent.
The moderator that tells you what to do. Where the human alone outperformed the AI, synergy was g = +0.46, p < 0.001. Where the AI alone outperformed the human, synergy was g = −0.54, p < 0.001, while human augmentation in that same condition was g = +0.74. Interaction F(1, 104) = 81.79, p < 0.001.
When you are worse than the model at a task, working with it makes you better than you were and makes the output worse than the model alone. You experience the improvement directly and the degradation not at all, which is why this level takes measurements instead of asking how it went.
Three qualifications, and one live example.
The example. Fifty physicians were randomised to GPT-4 access or conventional resources for diagnostic reasoning (Goh et al., JAMA Network Open, 2024). Physicians alone: median 74 per cent. Physicians with GPT-4: 76 per cent, adjusted difference 2 percentage points, 95 per cent CI [−4, 8], p = 0.60. GPT-4 alone: 92 per cent, beating the unassisted physicians by 16 points, p = 0.03.
Doctors given a tool that outperformed them by sixteen points improved by two, and not significantly. That is the meta-analytic finding happening in a hospital.
The heterogeneity is extreme. I² was 97.7 per cent for synergy. The pooled estimate is an average over wildly different studies, and individual settings depart from it substantially. Anyone quoting the −0.23 without this is overstating it.
The publication bias runs the unusual way. For synergy there was no evidence of bias, Egger's β = −0.67, p = 0.438. For human augmentation there was clear bias toward positive findings, Egger's β = 1.96, p = 0.002. The pessimistic headline is clean. The optimistic finding is the inflated one.
And there is a moderator that may rescue it. A 2026 re-analysis recoded all 370 conditions in the same corpus for whether participants received outcome feedback. Only 10 studies, 14 per cent, did. Studies with feedback showed synergy at g = 0.12; studies without at g = −0.17. Explanations plus feedback gave +0.30; explanations without feedback gave −0.31, with a Bayes factor of 21 for the moderation (Berger et al., 2026).
This does not overturn the meta-analysis; it identifies a moderator within the same data, observationally rather than experimentally. But it has a sharp practical consequence. Most professional work has delayed or absent outcome feedback. You write the memo and never learn whether the advice was right. If feedback is what makes the combination work, most of your working life is structurally in the g ≈ −0.17 regime.
Which is why Level 0 asked you to name a ground-truth task, and why writing "I do not have one" was a legitimate answer with consequences.
One caveat on the corpus that cuts the other way. The studies are from 2020 to 2023, so the systems are largely pre-GPT-4 classifiers rather than frontier models, and the corpus is dominated by decision tasks because those are the tasks where AI-alone performance can be computed at all. Creation tasks are under-represented for the same reason. That is a real threat to how far this transfers to generative work, and the authors say so.
Which statement about human-AI combination is supported?
Why does the outcome-feedback re-analysis matter for professional work?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.