Lead What Matters
A practicum in deciding what deserves your attention, acting on it under pressure, and proving over thirty days that the decision held.
- Levels
- 6
- Lessons
- 33
- Knowledge checks
- 87
- Gates
- 44
- Working tools
- 10
- Simulations
- 2
Most people entering leadership are not short of advice. They have been told to prioritise, to communicate clearly, to give feedback, to set boundaries and to learn from failure. The advice is broadly correct and it has not changed their Tuesday.
The gap is not informational. It is that every one of those behaviours has a cost that nobody names, and the cost arrives before the benefit. Prioritising means telling someone their work is not the priority. Feedback means saying the sentence you have avoided for six weeks. A boundary means a colleague is disappointed in you at 16:40 on a Thursday and you have to keep working next to them.
This course is built around those costs rather than around the benefits. Each level ends with something you have to do to a real person in a real organisation, and a set of gates a stranger could watch you fail.
You will finish with a written operating system for your own leadership, a thirty day record of whether you followed it, and at least one rule you had to change because it did not survive contact with a bad week.
What the evidence says about courses like this one
Leadership training works, on average, and the average conceals a great deal.
The largest meta-analysis covers 335 independent studies and 26,573 people. The corrected effect on behaviour transfer, meaning what leaders actually do afterwards, is 0.82, and the overall effect is 0.76 (Lacerenza et al., 2017). Those are substantial figures. Against them stands the same paper's own accounting of its evidence base: 208 of the 335 studies used a repeated-measures design with no control group at all, and roughly 81 per cent lacked a true control condition.
Across a century of leadership intervention studies, 200 in total, the average intervention moved the probability of a positive outcome from a coin flip to about 66 per cent (Avolio et al., 2009).
So the honest position is that this kind of course reliably produces something, that the something is smaller than the headline figures suggest, and that the design details matter more than the content. The same meta-analysis names them: transfer effects were far larger where a needs analysis was done first, where sessions were spaced rather than massed, and where feedback was built in. This course is spaced across sixteen weeks, opens with two weeks of measurement before it teaches anything, and requires you to collect feedback from named people in Level 4. Those are not stylistic choices.
Deliberate practice explains 14 per cent of performance variance across 88 studies and 11,135 people. Broken down by domain: games 24 per cent, music 21 to 23 per cent, sports 18 to 20 per cent, education 4 to 5 per cent, and professions 1 per cent (Macnamara, Hambrick & Oswald, 2014).
One per cent. In professional work, structured practice is a small term in a large equation. This course will make you better at a set of specific behaviours. It will not make you a leader, and anyone who tells you a programme can has not read the literature they cite.
The six plateaus
You will move through six capabilities, each gated by behaviour someone else could watch you fail.
First you measure where your attention actually goes, using a log rather than an impression. Then you build a filter that decides which concerns deserve weight, and attach your stated values to it as decision rules. Then you separate what you are responsible for from what you are being blamed for, and learn what that separation costs the people around you. Then you act against the pull of approval: the conversation you are avoiding, and the request you need to refuse. Then you extract information from failure and from feedback, including the badly delivered kind. Finally you write the rules you will run on, and test them for thirty days against a real workplace problem.
Cadence
Plan for three to four hours each week: about an hour of lessons and checks, two hours on the exercises against your own role, and twenty minutes on the weekly review. Levels 0 to 2 take two weeks each, Level 3 takes three because it contains two conversations you will want to postpone, Level 4 takes two, and Level 5 takes five because the final project is thirty days of dated evidence.
The source programme this course is built from runs eight weeks of facilitated sessions. This runs sixteen because a facilitated session can be attended and an exercise against a real colleague cannot.
Where the evidence comes from, and where it does not
Almost every psychological finding in this course was established on a Western sample. In an audit of the field, 96 per cent of research subjects came from Western industrialised countries holding 12 per cent of the world's population, and 68 per cent came from the United States alone (Henrich, Heine & Norenzayan, 2010). Eight years later, 94.15 per cent of samples in one flagship journal were still Western and 57.76 per cent still American (Rad, Martingano & Ginges, 2018). A 2026 audit of five British journals found 67.6 per cent of participants from Western countries, with 8.7 per cent drawn from Latin America, Eastern Europe, the Middle East and Africa combined (Petrutiu et al., 2026).
There is a second gap specific to this subject. A published audit of the geographic and occupational composition of leadership research samples was searched for and not found. The general figures above are the best available proxy, and a proxy is what they are.
Hierarchy, the acceptability of refusing a senior request, and what a boundary costs are all culturally loaded. The refusal script in Level 3 was validated nowhere, and the evidence that people overestimate the cost of declining was collected among American and Western European samples.
Run the scripts. Record what happened. Where your record and this course disagree, the record is the evidence and this course is the hypothesis.
What this course does not claim
It does not claim that a personal system fixes a structural problem. A team with three people doing the work of six will not be repaired by better boundaries, and a manager whose own manager punishes honest reporting cannot be trained into honest reporting.
It also does not claim indifference is the goal. The source material is precise about this and the phrasing is worth keeping: the course does not teach indifference, it teaches disciplined concern. A leader who cares about nothing is not steady, they are absent.
A level is passed when you have done something to a real person and can show what happened. Agreeing with a lesson costs nothing and produces nothing. This page has already changed nothing, and it is not supposed to.
The levels
- 0See where your attention actually goesReplace your account of how you spend your working day with a log that a neutral observer could have kept.Free
- 1Decide what deserves your concernBuild a filter that assigns weight to issues before you react to them, and attach your stated values to it as decision rules rather than descriptions.Locked
- 2Own the part that is yoursSeparate what you are responsible for from what you are being blamed for, take the first without the second, and learn what the separation costs the people around you.Locked
- 3Act against the pull of approvalHold the conversation you have been postponing, and refuse one request you would normally have absorbed, and record what each actually cost.Locked
- 4Use what goes wrongRun a review that names a wrong assumption rather than a person at fault, and collect feedback from three named people and sort it rather than swallowing or rejecting it.
Leadership ground
One real role, one real team, and one standing problem you carry through every exercise and into the thirty day project.
Before you finish Level 0 you fix three things. They stay fixed for sixteen weeks.
One: the ground
The role you actually hold now, with the people actually in it. Not the role you are being considered for and not the team you would have chosen. If you do not manage anyone, the ground is the set of working relationships where your judgement changes what happens: a project you coordinate, a client you handle, a founder you report to, three peers whose work depends on the order you do yours in.
Leadership without authority is the harder case and it is the common one. Every exercise in this course works without a reporting line. None of them work without named people.
Two: the standing problem
One problem in that ground that has been true for at least two months, that you have complained about at least once, and that you have the standing to address. It becomes the subject of the thirty day project in Level 5, and you will look at it from a different angle in each level before you touch it.
Examples that qualify: a recurring meeting that produces nothing and that you attend anyway. A colleague whose work arrives late and whose lateness nobody names. A reporting line so unclear that two people do the same task. A client who sends requests at 21:00 and receives replies at 21:04. A decision you have deferred four times, with dates.
Examples that do not: a problem entirely inside someone else's authority, where your only available action is to complain to a third party. A problem you have known about for one week, because there is no avoidance pattern to examine yet. A problem you would be unwilling to write down honestly.
If two problems come to mind and one of them involves a specific person you would have to speak to, that is the one with the pattern in it. The other one is a process complaint, and process complaints are where leadership courses go to die comfortably.
Three: the two named people
One person who will read what you write and tell you when it is evasive. One person inside the ground who will be affected by the thirty day project and who you will have to talk to about it.
Name both before Level 0 closes, in writing, with the date you asked them.
"Programme coordinator, health portfolio, four implementing partners and two analysts who report to me in practice but not on paper. The standing problem is the Thursday partner call: ninety minutes, eight attendees, no decisions since April, and I have complained about it to two people and changed nothing. My reader is Ada. The affected person is Ibrahim, who chairs it."
"I want to become a more confident and strategic leader who communicates better and manages my time well." No people, no dates, no problem, nothing that could fail.
Rules of practice
Seven rules. Four are supported by evidence, three are design choices, and the difference is marked.
1. Everything is written and dated
Design choice, with support by analogy. The evidence for recording comes from goal monitoring rather than leadership: prompting people to monitor progress raised attainment by d = 0.40 across 138 trials and 19,951 people, and physically recording progress outperformed tracking it without a record, 0.43 against 0.29 (Harkin et al., 2016). That literature is about personal goals, not about managing people. What survives the transfer is narrow and sufficient: an undated recollection of how a conversation went is an account written by the person with the strongest interest in the outcome.
2. Every exercise runs against a named person
Design choice. Not a scenario, not a composite, not "a stakeholder". The reason is in the evidence on transfer: leadership training moves behaviour most where practice and feedback are built in, and least where it is self-administered and abstract (Lacerenza et al., 2017).
3. You act before you feel ready
Evidence: moderate, by analogy. The strongest support is from a different field. Behavioural approaches that schedule and complete actions without first repairing mood or belief produce a standardised mean difference of 0.74 against controls across 26 trials in depression treatment (Ekers et al., 2014). Acting first is not borrowed from motivational writing. It is a treatment that matches the alternative which works on beliefs first.
4. Feedback is sorted, not swallowed and not rejected
Evidence: strong. Across 607 effect sizes and 12,652 participants, feedback interventions raised performance by an average d = 0.41, and more than 38 per cent of the effects were negative (Kluger & DeNisi, 1996). Feedback is not a good in itself. Level 4 is built on sorting it.
5. The record beats the framework
Evidence: strong. Where a specified action, taken for a specified period, produces no specified result, believe your record rather than this course. Two thirds of the honest literature on this subject reports smaller effects than the training industry built on it.
6. Nothing here requires anyone else to change first
Design choice. Every gate can be claimed inside a team that stays exactly as it is. That is a constraint on the course, not a claim about your organisation.
7. You claim your own gates
Design choice. No one marks these for you, and marking them early costs only you.
One lesson block read, the knowledge checks answered wrong at least once, the exercise run against a named person, three dated lines in the record, and one thing you avoided named honestly with the reason you avoided it.
All lessons read, all checks correct on the first attempt, the exercise done in your head about a hypothetical colleague, the record blank, and a private sense that you already do most of this.
Assessment
What counts, what does not, and what you should be able to show at the end.
What is assessed
| Component | Weight | The evidence |
|---|---|---|
| The attention log and its correction | 15% | Three days of timed entries, and what changed in the two weeks after |
| Exercise records | 20% | Dated entries against named people, one per exercise |
| Conversations held | 20% | The difficult conversation and the refusal, with what you feared and what happened |
| Feedback collected and sorted | 15% | Three named people's answers, recorded without rebuttal, and what you did with each |
| The thirty day project | 30% | The daily record, four weekly reviews, the result, and what failed |
What is not assessed
Confidence is not assessed. Fluency in discussing your leadership style is not assessed. Whether people like the changes you make is not assessed, and Level 2 explains why it cannot be.
A measurement warning that applies to your own self-assessment
Self-ratings are the least reliable instrument in this course, and there is a specific finding to prove it. Across 24 longitudinal studies of multisource feedback, the corrected improvement was 0.15 as rated by direct reports, 0.15 by supervisors, 0.05 by peers, and minus 0.04 by the leaders themselves (Smither, London & Reilly, 2005). Leaders rated themselves as having got slightly worse while everyone else recorded a small improvement, which tells you how weakly self-perception tracks anything.
This is why every exercise asks for dated records of specific acts rather than for ratings of your tendencies.
The check
Twelve statements, each scored one to five, run before Level 0 and again after Level 5. Run it now and record the date. The total is close to meaningless. The change on individual lines carries the information, particularly the lines on declining a request, holding a conversation you were avoiding, and acting without certainty.
What you should hold at the end
A three day attention log and a written account of what you changed because of it. A concern filter you have run on at least five live issues, with the dispositions recorded. Three values with one observable behaviour attached to each. A control map for your standing problem. One difficult conversation and one refusal, both dated, both with the outcome recorded including the parts that went badly. Feedback from three named people, sorted. A six part operating system. Thirty days of dated entries, four weekly reviews, and one rule you rewrote because a bad week broke it.
Sources
Every figure in this course, with the study behind it and what it can and cannot support.
Full citations with links are in SOURCES.md in the course directory. This page is the map: what each level rests on, and how strong it is.
How to read the grades
Strong means large samples, randomised or prospective designs, and either replication or a bias-corrected estimate. Moderate means well conducted but single-source, or meta-analytic with known publication bias. Weak means small samples, one laboratory, a commercial survey, or a claim that circulates without a controlled test behind it.
Where a figure is weak, the lesson says so rather than this page.
Level 0, on attention
| Finding | Figure | Grade |
|---|---|---|
| Office workers switch working spheres roughly every 11 minutes | 11 min 4 s mean; 57.1% of segments interrupted; N = 24 observed for 25 h 42 min each (Mark, González & Harris, 2005) | Moderate, small observational sample |
| Resuming an interrupted task the same day takes about 25 minutes | 25 min 26 s mean, 2.26 intervening spheres (Mark et al., 2005) | Moderate |
| Interruption makes people work faster and costs stress and effort | completion 22.77 min uninterrupted vs 20.31 and 20.60 interrupted; stress 6.92 vs 9.46 and 9.13, p < .01; N = 48 (Mark, Gudith & Klocke, 2008) | Strong for direction, weak for generalisation |
| Email is opened far more often than people believe | 74 email visits per day, mean visit 32.06 s; 566 application switches per day; N = 32 logged (Mark et al., 2015) | Moderate |
| Limiting email checking lowers daily stress | 1.46 vs 1.55, d = .37, p = .04; checks 4.70 vs 12.54 per day; N = 124 crossover (Kushlev & Dunn, 2015) | Moderate, small effect |
| "23 minutes 15 seconds to refocus" | appears in no peer-reviewed paper by the researcher it is attributed to | Debunked |
Level 1, on concern and values
| Finding | Figure | Grade |
|---|---|---|
| Hurry cut helping by a factor of six, in seminary students on their way to preach on the Good Samaritan | 63% low hurry, 45% intermediate, 10% high hurry; message content not significant; N = 40, F = 3.56, p < .05 (Darley & Batson, 1973) | Weak as evidence, strong as a case; never directly replicated |
| Stated values track behaviour unevenly across domains | strongest stimulation r = .68, tradition .67, hedonism .62; weakest security, conformity, achievement, benevolence r ≈ .30 to .39; three studies including partner-rated behaviour (Bardi & Schwartz, 2003) | Moderate |
| Moral licensing is real, small, and inflated by publication bias | d = 0.31, k = 91, N = 7,397; published 0.43 vs unpublished 0.11; three direct replications null, combined d = 0.07 (Blanken et al., 2015; 2014) | Strong as meta-analysis |
| Situational pressure to comply remains near 1960s levels | 70% continued past 150 volts, N = 40; 63% after seeing a refusal, N = 30; Milgram's comparable figure 82% (Burger, 2009) | Moderate, ethically constrained partial replication |
Level 2, on responsibility
| Finding | Figure | Grade |
|---|---|---|
| Believing you influence work outcomes tracks satisfaction strongly and performance weakly | satisfaction r = .34, k = 69, N = 16,348; performance r = .16, k = 9; burnout r = −.38, k = 5 (Wang, Bowling & Eschleman, 2010) | Strong for satisfaction, weak for the rest |
| Judging a situation controllable reduces sympathy and increases anger and neglect | 64 investigations, over 12,000 subjects; held across cultures, sample types and publication status (Rudolph et al., 2004) | Strong on direction |
| The same result under preregistration | N = 804 registered report; controllability pathway effects all η²p ≥ .03 (Yeung & Feldman, 2026, replicating Weiner, Perry & Magnusson, 1988) | Strong |
Level 3, on discomfort and refusal
| Finding | Figure | Grade |
|---|---|---|
| People underestimate how likely others are to agree to a direct request | participants overestimated the number of people needed by roughly 50%; volunteers also underestimated average donation by $17 (Flynn & Bohns, 2008) | Moderate, author-reported magnitude |
| A request made in person vastly outperforms the same request by email | roughly 8 in 10 in person against a fraction of 1 in 10 by email; 45 participants, 450 targets (Roghanizad & Bohns, 2017) | Moderate |
| People overestimate how much declining upsets the person who asked | 5 experiments, N = 406 and 208 couples; direction consistent, effect sizes not obtained (Givi & Kirk, 2024) | Moderate, direction only |
| Senders withhold and soften bad news, reliably, since 1970 | the MUM effect; replicated across laboratory and field, no clean percentage available (Rosen & Tesser, 1970; Dibble & Levine, 2010) | Moderate, teach the effect without a number |
| About a quarter of participants acting as managers withheld performance information even when it was positive | N = 2,620 randomised online experiment; withholding nearly twice as likely toward women when negative information was imprecise (Management Science, 2026) | Moderate, role-play design |
Level 4, on failure and feedback
| Finding | Figure | Grade |
|---|---|---|
| Feedback raises performance on average and harms it often | d = 0.41; over 38% of 607 effects negative; N = 12,652 (Kluger & DeNisi, 1996) | Strong |
| Information-carrying feedback beats reinforcing feedback by about four to one | d = 0.99 against 0.24; 32 meta-analyses, 435 studies, over 61,000 participants (Wisniewski, Zierer & Hattie, 2020) | Strong, educational settings |
| Multisource feedback produces improvement below the threshold for "small" | corrected d: direct reports .15, supervisors .15, peers .05, self −.04; 24 longitudinal studies (Smither, London & Reilly, 2005) | Strong |
| People want the hard feedback more than the person holding it believes | 2.6% told a stranger about a visible mark on their face; 5 experiments, N = 1,984 (Abi-Esber et al., 2022) | Moderate, and see the integrity note in SOURCES.md |
| Negative feedback works only where trust exists and expectations were set in advance | 173 studies synthesised (Heine, Liden & Stouten, 2026) | Moderate, narrative synthesis |
Level 5, on systems and the thirty day project
| Finding | Figure | Grade |
|---|---|---|
| Time to automaticity is months and varies enormously | median 66 days, range 18 to 254, among 39 of 96 recruited; exercise median 91 days (Lally et al., 2010) | Moderate |
| A 2024 review confirms the range and the variability | medians 59 to 66 days, means 106 to 154, individual range 4 to 335; 20 studies, N = 2,601 (Singh et al., 2024) | Strong |
| Missing one day does not break a habit | automaticity fell 0.29 points, difference non-significant (Lally et al., 2010) | Moderate |
| Leadership training transfers best when spaced, when feedback is included, and when a needs analysis came first | transfer δ .92 spaced against .45 massed; 1.40 with feedback against .50 without; 335 studies, N = 26,573 (Lacerenza et al., 2017) | Moderate for the subgroup values |
| "21 days to form a habit" | originates in a 1960 book by a plastic surgeon describing how long patients took to stop seeing their old face | Debunked |
Claims removed from this course, and why
"It takes 23 minutes and 15 seconds to refocus after an interruption." Traced to a 2006 interview remark and present in none of the researcher's papers. The published figures are 25 min 26 s for same-day resumption and 20.31 to 22.77 minutes for task completion. Replaced by those.
"93 per cent of communication is nonverbal." Assembled from two 1967 studies that never compared the three channels simultaneously, one of which used 37 female psychology majors and a single spoken word. The originator called the generalisation absurd.
"You need a 3:1 ratio of positive to negative interactions." The mathematics behind the 2.9013 threshold was shown to be borrowed from fluid dynamics and fundamentally in error, and the senior author publicly withdrew the modelling.
"10,000 hours of deliberate practice makes an expert." Deliberate practice explains 1 per cent of performance variance in the professions.
"Match your delivery to people's learning styles." No adequate evidence base; the studies with a design capable of testing it contradict it.
"360 degree feedback transforms leaders." Corrected improvement of 0.15 at best, and minus 0.04 on self-ratings.
"Managers who avoid difficult conversations cost their organisations X." Every figure located for this was a non-probability commercial panel survey commissioned by a firm selling communication training. No study was found that measured avoidance in real organisations with a cost attached. The course therefore teaches the behaviour without claiming a price for it.
Maslow's pyramid. Maslow did not draw a pyramid. It was constructed later by management writers.
What is open, and what is not
The free levels of this programme are readable with no account at all — the real levels, not samples. An account carries your progress, your gate claims and your saved work. The remaining levels, the tools, the gates and this programme’s worked exemplars are opened together when you enrol.
Create an account