Why your own judgement will not do
The evidence in this lesson is about people rather than models, which is why it is likely to outlast the models.
The perception gap. The METR developers give the cleanest version. Four separate predictions, all wrong in the same direction. Participating developers before the study: 24 per cent speedup. The same developers after completing the tasks: 20 per cent speedup. Economics experts: 39 per cent. Machine-learning experts: 38 per cent. Measured: 19 per cent slowdown.
The participants had just done the work and had screen recordings of themselves doing it.
What trust in the tool does to attention. A survey of 319 knowledge workers, all using generative AI at work at least weekly, produced 936 real workplace examples. Between 55 and 79 per cent reported less cognitive effort across every activity type, with comprehension at 79 per cent and evaluation at 55 per cent. Over 40 per cent of real uses involved no critical thinking the user could identify.
Two coefficients matter, and they run against each other. Confidence in the AI was negatively associated with critical thinking, β = −0.69, p < 0.001. Confidence in one's own expertise was positively associated with it, β = 0.26, p = 0.026 (Lee, Sarkar, Tankelevitch et al., CHI 2025).
The more you trust the tool the less you check. The more you trust yourself the more you check. Both hold at once, which is why this course spends as much effort keeping your own judgement in play as it spends on checking the model.
Hold this study at its actual strength, which the authors are careful about. It is entirely self-report. It measures how knowledge workers experience their own thinking, with no objective performance measure. Given the 39-point self-report error above, that is a serious limit, and the secondary coverage of this paper routinely ignores it.
The experimental evidence arrived later and is stronger.
Three randomised experiments with 1,222 participants had people work through learning problems with or without AI, then tested them with no AI for anyone (Liu, Christian, Dumbalska, Bakker & Dubey, 2026).
Experiment 1, fractions: test solve rate 0.57 with AI against 0.73 without, t(305) = −3.64, p < 0.001, d = −0.42. Experiment 2, with pretest controls and equal exclusions: 0.71 against 0.77, t(583) = −2.33, p = 0.020, d = −0.19. Experiment 3, reading comprehension: 0.76 against 0.89, t(166) = −2.72, p = 0.007, d = −0.42.
The dose-response inside the AI condition is the useful part, and it is cross-sectional rather than randomised, so read it as a pattern rather than a finding. Participants who requested direct answers scored 0.65. Those who requested hints only scored 0.76. Those who had access and did not use it scored 0.89. Control scored 0.77.
Hint users were indistinguishable from the control group. Answer users were significantly worse, t(464) = −3.77, p < 0.001, d = −0.36. This is the closest thing in the literature to an evidence-based instruction about how to use the thing, and it emerged after roughly ten minutes of exposure.
A second randomised experiment, 52 experienced Python developers learning an unfamiliar library over about an hour, found the AI condition scoring 17 per cent lower on a 27-point assessment, d = 0.738, p = 0.010, with the largest gaps in debugging, and no significant productivity gain (Shen & Tamkin, 2026). No trade-off to weigh. A straight loss. Note N = 52, which makes d = 0.738 an upper bound pending replication.
Both studies measure failure to acquire a new skill over minutes or hours. Neither measures decay of an existing professional skill over months, which is the thing people actually worry about. That remains unmeasured, and Level 5 says so again.
What is the correct strength to assign the critical-thinking survey?
What does the dose-response pattern in the learning experiments support?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.