The jagged frontier
The single most important study for this course built two tasks on purpose: one the researchers believed sat inside the model's competence, one outside it, run on the same population with the same design.
N = 758 consultants at a global firm, about seven per cent of its individual-contributor workforce, in two preregistered experiments with three arms each: no AI, GPT-4, and GPT-4 plus a prompt-engineering overview.
Inside the frontier. A creative product-innovation task across eighteen realistic subtasks. Quality rose 33.9 per cent with the overview and 29.9 per cent without it, in the published version. Consultants completed 12.2 per cent more subtasks and worked 25.1 per cent faster. All p < 0.01.
Note that the "about 40 per cent" figure in circulation comes from the 2023 working paper, which reported 42.5 and 38 per cent. The peer-reviewed version reports 33.9 and 29.9. Use the published one.
Outside the frontier. A business problem requiring reconciliation of spreadsheet data with interview material, constructed so that the obvious statistical reading was wrong and the correct answer required noticing the conflict.
Control: 84.5 per cent correct. GPT-4 plus overview: −24.5 percentage points. GPT-4 alone: −13.9 percentage points. Averaged, a 19 percentage point drop, all p < 0.01. The error rate went from about 15.5 per cent to about 35 per cent.
And on that same task the rated quality of recommendations rose, by 25.1 per cent with the overview and 17.9 per cent without, including among consultants who gave the wrong answer. They were also 18 to 30 per cent faster.
Nothing. On the task where the error rate more than doubled, the work came back faster, better written and more confident, and its rated quality rose, including for the consultants who were wrong. Every signal you would use to catch the failure improved with the failure. If you are waiting for AI-assisted work to feel off before you check it, you are waiting for a signal that moves the wrong way.
Three things follow.
The frontier is jagged and you cannot see its shape. The two tasks were both plausible consulting work. Nothing on the surface distinguished the one where AI added a third to quality from the one where it doubled the error rate. That is the whole meaning of the word jagged: capability does not decline smoothly with difficulty, so a task you find hard may sit inside and a task you find easy may sit outside.
Training in prompting made the outside-frontier result worse, not better. The arm that received the prompt-engineering overview lost 24.5 points against 13.9 for the arm with no training at all. The same appendix shows why: the modal retention score was about 0.87, meaning most consultants kept nearly all of what the model produced, and trained subjects showed more extreme retention rather than less. Prompt training increased copy-pasting.
And it does not level the field everywhere. Inside the frontier, bottom-half performers improved 43 per cent against 17 per cent for the top half, which is the levelling story everyone tells. Outside it, the pattern inverts and disparities amplify. A field experiment with 640 Kenyan small-business owners found the same inversion at the population level: no average effect, with low performers at −8 per cent and high performers at +15 per cent, difference p = 0.013 (Otis et al., 2025). AI levels the playing field when the population's hard problems fall inside the frontier. When they fall outside it, it does the opposite.
One honest note about a pair of terms you will meet. The same paper describes two working styles observed in interaction logs: Centaurs, who split work at a clean boundary, and Cyborgs, who interleave continuously. These are descriptions, drawn from logs, and the authors say they "require deeper examination". No trial has trained anyone into either pattern and measured the result. This course does not recommend them.
Why is the outside-frontier result more important than the inside-frontier result?
A colleague says the consulting study shows AI levels the playing field between weak and strong performers. What is the accurate correction?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.