The three counts
Three numbers, from one piece of your own work, taken now and again in week eight. Everything in this course is judged against the change.
Count one: retention. For one AI-assisted output, what proportion of the model's text survived into what you sent, unchanged. Count words if you can. Otherwise count sentences: how many of the sentences you sent are word-for-word what came back.
The reference points. Sixty-eight per cent of participants in the writing experiment submitted output unedited. The modal retention score among consultants was about 0.87. These are not targets and they are not failures. They are the context in which your own number is either ordinary or remarkable, and it will almost certainly be ordinary.
Count two: the judge's four answers. Give the finished piece to your named judge, without telling them AI was involved, and ask four questions.
"What in here would you act on?"
"What in here would you want checked before acting?"
"What is missing?"
Write the answers verbatim before you respond to any of them. Then write your own four answers, separately, and keep both.
Count three: verified claims. How many factual claims in the piece did you check against a source that was not the model. Not "how many looked right". How many you opened something to confirm. Record the number and the total number of checkable claims in the piece.
For most people, in most weeks, this number is zero, and writing the zero down is the exercise.
Two design notes, because both counts are deliberately cruder than they could be.
Retention is a proportion rather than a rating. You are not scoring how much you improved the draft. You are counting what survived. The reason is that "how much did I improve it" is a judgement, and this level exists because judgements about your own AI-assisted work have been measured and found wrong by about 39 percentage points in the flattering direction. A proportion is a thing two people would count the same way, which is what makes week one and week eight comparable.
The judge does not know which piece is which. In week eight you will give them a second piece and ask the same four questions. If they know which is the "after" one, you have a rating of your progress rather than a measurement of your work. This matters more than it sounds: the meta-analytic evidence on explanations in human-AI decision work found that explanations increased acceptance of AI recommendations unconditionally, improving accuracy when the recommendation was right and worsening it when it was wrong (Bansal et al., CHI 2021). Telling somebody what to expect moves their judgement without improving it.
And a warning about the third count. Verification is where this course spends Level 4, and the reason it is measured from week one is that nobody knows what it costs. A systematic search for any empirical estimate of verification time as a proportion of task time returns nothing in any domain. The nearest figure that exists anywhere is METR's screen-recording breakdown, where developers spent about 9 per cent of their total time reviewing and cleaning AI output. That is one number, from one study, in one domain, and it is the best the field has. Your own log will be better evidence about your own work than anything published.
Why does this level ask for a retention proportion rather than a rating of how much you improved the draft?
Your verification count for the week comes out at zero. What should you do?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.
This lesson has a tool
Open it and get your draft reviewed. Drag-and-drop tools need a wider screen; the review works anywhere.