The four that survive
If the catalogue is out, something has to replace it. Here is the replacement: four patterns with strong replication records, chosen because each one has an identifiable moment in a working week where it operates, and each has a written control that fits inside that moment.
Anchoring: the first number moves the rest. In a multi-lab replication with 6,344 participants across 36 samples, standard anchoring returned effect sizes between 1.27 and 2.60, larger than in the original demonstration (Klein et al., 2014). It operates on professionals with the tools of their trade. Across four studies, legal professionals gave higher sentences after a higher number was put in front of them, including when a journalist supplied it in a phone call and when the participant generated it by throwing dice, with the paper concluding that expertise and experience did not reduce the effect (Englich, Mussweiler and Strack, 2006; Ns of 42, 39, 52 and 57). Note what the classroom version of that study usually gets wrong: the dice condition used postgraduate legal trainees with a mean age of 27.5, not sitting judges. The two studies that used experienced judges used a journalist's question and a prosecutor's demand, which is if anything closer to your working week.
Moment: any meeting where a figure is proposed before the estimate is made. Control: write your own number before the meeting, and if you cannot, say that you will send it afterwards.
Sunk cost and escalation: the money already spent buys more of it. Sunk cost replicated cleanly at d = 0.31 in the same multi-lab project, and in an incentivised design where the inferior option was genuinely dominated, 23 per cent of participants who had earned it kept it, against 7 per cent who were simply given it and none of those who chose freely (Ronayne, Sgroi and Tuckwell, 2021; N = 528). The meta-analysis of escalation is the more useful document for professional life. Across 166 independent samples the strongest correlate of escalation was not sunk cost at all. It was proximity to project completion, at ρ = .393 against .243 for sunk costs and .258 for personal responsibility for the original decision (Sleesman et al., 2012).
Moment: the project at 80 per cent. Control: the question is not what has been spent, it is what you would commit today to reach the current position from a standing start.
Biased assimilation: the evidence you agree with is better evidence. Presented with two studies on a contested question, people rate the one supporting their prior as more convincing and better conducted. This survives at scale: in three experiments with about 5,632 participants using arguments personalised to each person's own core issue, biased assimilation replicated with gaps of 1.22 to 2.69 scale points, and counter-attitudinal arguments drew 16 percentage points more denigrating responses (Velez and Liu, 2025). What did not survive is the more dramatic half of the original claim. Attitude polarisation, the finding that mixed evidence pushes people further apart, was indistinguishable from zero in two of those three experiments, and the third found movement towards moderation.
Moment: reading two sources that disagree. Control: write your rating of methodological quality before you read the conclusions, or, if that is impractical, write the two strongest objections to the source you agree with before writing anything about the other.
Overprecision: your ranges are too narrow. Of the three things usually bundled as overconfidence, only this one is robust. Overestimation of one's own performance is described in the review literature as thin and inconsistent; overplacement reverses on hard tasks, where people place themselves below average. Overprecision persists. In the illustrative study, intervals offered as 90 per cent confident contained the true answer 73.1 per cent of the time, across all difficulty levels (Moore and Healy, 2008; N = 82), and the review reports confidence-interval studies with hit rates below 50 per cent (Moore and Schatz, 2017).
Moment: every estimate, forecast and timeline you give. Control: widen the range until it feels slightly embarrassing, then check it later. Level 3 makes this measurable on your own numbers.
The list is short on purpose. A control you apply at three identifiable points in a week is worth more than a catalogue you can recite and never deploy. If you want a fifth, take it from your own decision records in Level 6 rather than from a textbook, because the one that costs you most is specific to the work you do.
Some things are missing from that list that you would find in most courses, and the omissions are deliberate.
Availability is one of the most-taught heuristics and has one of the thinnest modern replication records. A direct replication of the famous-names demonstration with 195 participants found no effect on raw estimates, and an attenuated effect of d = 0.34 only on a derived measure (McKelvie and Drumheller, 2001). It is not covered by either large multi-lab project. The underlying idea, that ease of recall stands in for frequency, remains plausible. The evidence for it is not in the same class as the four above.
Base-rate neglect in its strong form was rejected by the critical literature thirty years ago. A review of the field concluded that base rates are almost always used and that the degree of use depends on task structure and how the problem is presented (Koehler, 1996). What survives is underweighting, conditional on format, and Level 1 covers the format that fixes most of it.
Vividness distorting probability did not merely fail to replicate. In a 125-sample project it reversed direction, from an original d = 0.74 to a replication d of −0.08 (Klein et al., 2018).
Groupthink is not on the list because it is not an evidenced construct, which Level 4 takes up in full.
There is a broader habit to take from this. When you next meet a confident claim about how people think, the useful question is not whether it sounds right. It is which of these it is: a finding that survived a large multi-lab replication, a single study from one team, or a compelling narrative that has been repeated until it acquired the texture of a finding. The three are indistinguishable in a conference talk and about equally common.
A capital project is 85 per cent complete, has overrun by nine months, and the business case no longer holds. The steering group votes to continue. Which factor does the meta-analytic evidence identify as the strongest correlate of that decision?
You are about to give a client a delivery estimate. Which of the four patterns is most likely to be operating, and what is the control?
Which claim about thinking errors is best supported by large-scale replication?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.