Make Sense of AI
A practicum in recognising the artificial intelligence already deciding things around you, using a chatbot on real work, checking what it tells you, and keeping the decisions that should stay yours.
- Levels
- 6
- Lessons
- 34
- Knowledge checks
- 83
- Gates
- 50
- Working tools
- 10
- Simulations
- 2
- Worked exemplars
- 4
Make Sense of AI is a practical introduction for people unfamiliar with or unsure about artificial intelligence. It helps you recognise the AI already around you and use it effectively and efficiently on real tasks.
This course starts with one real task and a named reader. AI can help you draft, explain, summarise, translate, organise options and explore a problem; dependable use means knowing what you want, checking what matters and keeping decisions that require your judgement.
You will build tangible artefacts as you go: an inventory of systems already in your week, a kept record of generated answers on a subject you know, a real request and its revisions, five verification records, a privacy list, an impersonation control and a six-part plan tested for thirty days.
Whether AI is new to you or already part of your week, this course offers a practical method. No programming, mathematics or prior use of a particular provider is required. Bring a phone or computer, a real task, a subject you know well and a person who can read your work.
The six plateaus
You move through six capabilities, each demonstrated by an artefact or action another person could inspect.
First you find the artificial intelligence already in your week, from a written inventory rather than an impression, and learn to tell a system that follows rules from one that learns patterns. Then you read a generated answer for what it actually is, and catch fabrication on a subject you know well. Then you instruct a chatbot well enough to finish a real task, and discover from your own record whether the disappointing answer was the tool or the request. Then you build a verification habit that survives being in a hurry. Then you decide what you will never type, and install one control against impersonation in your family and one in your work. Finally you write the rules you will run on and test them for thirty days against a real problem.
Cadence
Plan for two to three hours each week: about forty minutes of lessons and checks, ninety minutes on the exercises against your own subject, and fifteen minutes on the record. Levels 0 to 4 take two weeks each. Level 5 takes five, because the final project is thirty days of dated entries and four weekly reviews.
The source form of this course is eight facilitated sessions of two hours. The fifteen-week self-paced form is a practice extension: a session can be attended, but a thirty day record cannot.
Why this structure works
The sequence pairs explanation with observable practice. In a controlled study of practice-oriented AI literacy, learners who combined explanation with drills accepted underspecified prompts less often, 51.5 per cent against 66.7 per cent, and asked follow-up questions more often, 59.2 per cent against 27.9 per cent, p < .001; the performance gain was small, r = .19 (Clerc et al., 2026). That convergence is why each level here produces an artefact rather than only an opinion.
The wider intervention evidence sets boundaries. A short animation changed knowledge, trust and intention in 592 Australian adults, but intentions may not translate directly into behaviour (Ayre et al., 2025). An explanation-only intervention left adoption of wrong recommendations unchanged at 52.1 per cent and increased rejection of correct suggestions (Puppart and Aru, 2025). An hour of in-person media-literacy training in India produced no average improvement and varied by political alignment (Badrinathan, 2021). Tips for spotting false news improved discernment in a representative United States sample, less in an educated urban Indian sample and not in a largely rural northern-India sample (Guess et al., 2020). These findings favour practice, context and explicit checks over a single persuasive lesson.
Self-report is useful as a reflection prompt, not as completion evidence. In the practice study, self-reported competence predicted measured performance at r = 0.01, p = .88. The course therefore asks for dated actions, source trails and revisions rather than a confidence score.
Spaced retesting gives the cadence its shape. Inoculation effects retained 100 per cent over three months when participants were re-tested at one, five and thirteen weeks, compared with 36 per cent without intermediate testing and no statistical significance after two months, F = 2.17, p = .14 (Maertens et al., 2021). The fifteen-week sequence, written record at each level and two post-course re-tests apply that finding while leaving room for the learner's own evidence.
Where the evidence comes from, and where it does not
Almost every figure in this course was collected outside Africa, and several were collected only in the United States.
In an audit of six psychology journals from 2003 to 2007, 96 per cent of research subjects came from Western industrialised countries holding 12 per cent of the world's population, and 68 per cent came from the United States alone (Arnett, quoted in Henrich, Heine and Norenzayan, 2010). The follow-up audit of the same journals from 2014 to 2018 found samples and authors still a little over 60 per cent American, and concluded that 89 per cent of the world's population continues to be neglected (Thalmayer, Toscanelli and Arnett, 2021).
The narrowing gets worse as the subject gets closer to this one. Across 2,768 papers at the main human-computer interaction conference from 2016 to 2020, 73.13 per cent of participant samples came from Western countries and 45.82 per cent from the United States alone. Of 195 countries, 102 contributed no participant samples at all, with the gaps concentrated especially in Africa (Linxen et al., 2021). At the flagship conference on fairness and accountability in artificial intelligence, 84 per cent of papers with human participants used participants exclusively from Western countries, and 93.6 per cent of all participants were American (Septiandri et al., 2023).
Three specific absences matter more than the general one.
No study measures whether people can identify a cloned voice speaking Nigerian English, Pidgin, Yoruba, Hausa or Igbo. Every detection figure in Level 4 comes from listeners in the United Kingdom, United States, Germany, China, Spain and Japan.
No documented case of artificial intelligence voice-cloning fraud in Nigeria was located: no police report, no charge, no court record, no named victim account. Nigeria's documented artificial intelligence fraud record is entirely video and image impersonation of public figures in investment schemes. Given how much of Nigerian communication runs through voice notes, that is far more likely to be an absence of documentation than an absence of crime, and this course will not assert either.
No survey measures how many Nigerian employers have a written policy on using these tools. Nigeria is included in the largest global survey of attitudes to artificial intelligence, but the published country-level result on that question is not available.
Run the exercises. Record what happens. Where your record and this course disagree, the record is the evidence and this course is the hypothesis.
There is a fourth problem, and it is the most interesting one. Three credible sources describe three different Nigerias. An online research panel of about a thousand Nigerians found 79 per cent willing to trust artificial intelligence and 73 per cent reporting some formal or informal training in it (Gillespie et al., 2025). A face-to-face survey found 17 per cent of Nigerians had heard or read a lot about artificial intelligence (Pew Research Center, 2025). Behavioural data from one major chatbot put Nigerian usage at 0.2 times the level expected for its working-age population, the lowest of any country measured (Anthropic, 2025).
None of these is wrong. They sample three different populations: connected urban Nigerians, Nigerians, and the customers of one product with almost no consumer presence in the country. Which one you belong to changes what this course is for.
What this course does not claim
It does not claim that these tools are unreliable and should be avoided. Across 106 experimental studies and 370 effect sizes, artificial intelligence assistance improved on humans working alone by a pooled Hedges' g of 0.64 (Vaccaro, Almaatouq and Malone, 2024). The gains are real and they are measured.
It does not claim that a chatbot is a substitute for a professional, a colleague or a document.
It does not claim that being careful is the same as being useful. A person who verifies everything and produces nothing has not learned this subject either.
It does not promise a particular attitude or automatic behaviour change. It gives you a structured opportunity to practise, test and revise, with a record that can show what happened in your setting. The evidence portfolio is the practical outcome, and it remains open to refinement as you learn from your own task.
A level is complete when you have done something with a real tool on a real subject and can show what happened, including the parts that needed correction. The record, not agreement with a lesson, is what lets you and your reader decide what to practise next.
The levels
- 0Find the systems already deciding for youReplace the picture of artificial intelligence as a future robot with a written inventory of the systems that sorted, ranked, scored or predicted something about you this week.Free
- 1Read a generated answer for what it isCatch a chatbot fabricating on a subject you know better than it does, keep the evidence, and stop using fluency as a signal.Locked
- 2Instruct it so the answer is usableTake one real task from a weak first request to a finished answer, and find out from your own record whether the disappointment was the tool or the instruction.Locked
- 3Check it before you act on itBuild a verification habit that survives being in a hurry, on claims that would cost you something if they were wrong.Locked
- 4Decide what you will not type, and install one controlRead what one provider currently does with what you type, write the list of what you will never enter, and install a verification control with the people who would call you about money.
Your AI ground
One real task, one real subject you already know well, and one person who will read what you write and tell you when you are wrong.
Before you finish Level 0 you fix four things. They stay fixed for fifteen weeks.
One: the task
One real thing you have to produce, in the next three months, that involves writing, explaining, summarising, planning or deciding. It must have a reader who is not you.
Examples that qualify. A funding proposal for a cooperative you belong to. A lesson plan for a class you actually teach. A letter to a landlord about a repair. A business plan for a shop you actually run. A study plan for an examination you are actually sitting. A report your supervisor is actually waiting for. A community event you are actually organising.
Examples that do not. A task you might do one day. A task with no reader. A task you would not be willing to show anyone.
Two: the subject you already know well
One subject where you can tell, without checking, whether a statement about it is true. Your own trade. Your own town. Your own family history. The rules of a sport you played. The regulations of the office you work in. The market prices you buy at.
This is the instrument you will use in Level 1 to catch the tool fabricating, and it only works if your knowledge is genuinely better than a plausible guess. If you cannot immediately say what would be wrong about a statement in your chosen subject, choose another.
Fabrication rates are not uniform. In a 2025 study of 176 citations generated across six literature reviews, entirely fabricated citations ran at 6 per cent for the most common of three disorders and 28 to 29 per cent for the two less common ones (Linardon et al., 2025).
The rate falls as the subject gets more common and rises as it gets more specific. Your local knowledge is therefore the sharpest available test, and the reason this course does not ask you to check the tool on general knowledge is that general knowledge is where it performs best.
Three: the standing exposure
One place where you or someone near you already gives information to a system without knowing what happens to it. A loan application on a phone. A bank verification process. A workplace tool that arrived without an explanation. A group where forwarded messages are believed. An elderly relative who answers unknown numbers.
It becomes the subject of Level 4 and of one control you install for real.
Four: the two named people
One person who will read what you write and tell you when it is evasive or wrong. One person who has never used a chatbot, who you will teach in Level 0 and again in Level 2.
The second person is not decoration. Explaining a thing to someone who has no idea what it is will find the parts you do not understand faster than any check in this course.
"The task is the annual report for the parents' association, due 12 November, read by about forty parents and the school head. The subject I know well is the transport route between Ikeja and Yaba: fares, hold-ups, which junctions flood. The standing exposure is my mother, who answers every number and forwards every message in three WhatsApp groups. My reader is Chidi. The person I am teaching is my mother."
"I want to get better at using AI and understand how it works so I can be more productive and keep up with technology." No task, no reader, no date, no subject anyone could be wrong about, nothing that could fail.
Rules of practice
Seven rules. Three are supported by evidence, four are design choices, and the difference is marked.
1. Every exercise runs against a real task with a real reader
Design choice. Not a scenario and not a practice prompt. The reason is in the evidence on transfer: the one artificial intelligence literacy intervention that moved behaviour rather than only knowledge was built on practice drills with real tasks, and the two that failed were explanation delivered once (Clerc et al., 2026; Puppart and Aru, 2025; Badrinathan, 2021).
2. Everything is written and dated
Design choice, with support by analogy. The evidence for recording comes from goal monitoring rather than from this subject: prompting people to monitor progress raised attainment by d = 0.40 across 138 trials and 19,951 people, and physically recording progress outperformed tracking it without a record, 0.43 against 0.29 (Harkin et al., 2016). What transfers is narrow and sufficient. An undated recollection of how a conversation with a chatbot went is an account written by the only party with an interest in the outcome.
3. You keep the whole conversation, including the bad answers
Design choice. A common shortcut is to delete exchanges where the tool was wrong and keep the useful ones, producing a record that agrees with the impression already held. Keep the wrong answers: they are the data.
4. A stated confidence is not evidence
Evidence: strong. When asked to state a confidence figure, models produce numbers clustered between 80 and 100 per cent, usually in multiples of five, and the measured accuracy in each band sits well below the figure stated (Xiong et al., ICLR 2024). Across eight artificial intelligence search tools answering 1,600 news queries, one tool misidentified 134 of 200 articles and signalled low confidence 15 times in 200 responses (Tow Center, 2025).
5. You check before you act, not before you believe
Evidence: strong, and it cuts against you. Across 607 effect sizes, help from a language model raised fact-checkers' accuracy from 59 per cent to 74 per cent on average, and dropped it to 35 per cent when the model's explanation happened to be wrong, below the 49 per cent they achieved with no help at all (Si et al., NAACL 2024). Assistance raises the average and lowers the floor. Verification is what keeps you off the floor.
6. You do not enter another person's information into a service you cannot describe
Design choice, with a legal edge in Nigeria. Under the General Application and Implementation Directive 2025, a data processor is expected to rely on a Data Processing Agreement with the data controller in order to carry out data processing. An employee typing a customer's file into a consumer chatbot in a browser has no such agreement.
7. You claim your own gates
Design choice. You claim these gates for yourself, using the artefacts and dates so a reader can inspect the same evidence.
One lesson block read, the knowledge checks answered wrong at least once, the exercise run on your real task, the whole conversation kept including the parts where the tool was wrong, three dated lines in the record, and one thing you did not check named honestly with the reason.
All lessons read, all checks correct on the first attempt, the exercise done in your head about a task you might do later, no conversation kept, and no record another person can inspect.
Assessment
What counts, what does not, and what you should be able to show at the end.
Evidence portfolio and completion standards
The percentages below are planning weights for a facilitator or learner deciding where to spend time; they are not a claim that arithmetic alone grades operational judgement. Completion is demonstrated by the observable evidence in the final column. A learner who has the evidence can explain the decisions, revisions and limits in it; a learner who has only a score has not yet shown the practice.
| Component | Weight | The evidence |
|---|---|---|
| The inventory and the recognition test | 15% | A week of logged systems, and one other person's answers to the six-item test |
| The fabrication record | 20% | Ten questions on your own subject, checked, with the wrong answers kept |
| The real task, start to finish | 20% | Every request and response, the time you spent, and what you had to fix |
| Verification records | 15% | Five claims run through a written check, each with the source and the disposition |
| Privacy and impersonation controls | 10% | The list of what you will not type, and one control installed and tested |
| The thirty day project | 20% | Daily entries, four weekly reviews, the result, and the rule that broke |
What is not assessed
Confidence is not assessed. Fluency in discussing artificial intelligence is not assessed. Whether you end the course enthusiastic or wary is not assessed, and Level 5 explains why neither is the goal. No external assessor is implied: the portfolio is a self-directed record, optionally reviewed by the named reader or facilitator.
A measurement warning that applies to your own self-assessment
Self-rating is the weakest instrument in this course, and there are two findings to prove it.
Students' self-reported competence with generative tools predicted nothing about their measured performance, r = 0.01, p = .88, and their self-reported metacognition predicted nothing either, r = 0.04, p = .65 (Clerc et al., 2026).
Sixteen experienced developers, working on real issues in repositories they already knew, forecast that artificial intelligence tools would make them 24 per cent faster. Afterwards they reported having been 20 per cent faster. Measurement showed them 19 per cent slower (METR, 2025). Their sense of their own speed was wrong by about forty percentage points, in the direction that flattered the tool.
This is why every exercise asks for a dated record of a specific act rather than a rating of your tendencies.
The check
Ten statements, each scored one to five, run before Level 0 and again after Level 5. Run it now and record the date. The total carries almost no information. The change on individual lines carries it, particularly the lines on checking claims, on knowing what happens to what you type, and on the difference between a confident answer and a correct one.
What you should hold at the end
A week of logged systems with what each one decided. One other person's answers to the six-item recognition test. Ten checked questions on a subject you know well, with the fabrications kept. A real task completed with the whole conversation, the time it took, and the list of what you fixed. Five verification records with sources. A written list of what you will not type, in your own words with your own examples. One impersonation control agreed with real people and tested once on a real call. A six-part use plan. Thirty days of dated entries, four weekly reviews, and one rule you rewrote because a bad week broke it.
Sources
Every figure in this course, with the study behind it and what it can and cannot support.
Full citations with links are in SOURCES.md in the course directory. This page is the map: what each level rests on, and how strong it is.
How to read the grades
Strong means large samples, randomised or prospective designs or official test data, and either replication or a published margin of error. Moderate means well conducted but single-source, or a preprint that has not been through peer review, or a regulator statement without figures attached. Weak means small samples, one laboratory, a vendor survey, or a claim that circulates without any controlled test behind it.
Where a figure is weak, the lesson says so rather than this page.
Level 0, on recognising the systems
Strong. The recognition gap: 99 per cent used at least one of six artificial intelligence products in the previous week, 36 per cent said yes when asked directly, 3,975 US adults, margin of error 2.6 points (Gallup, 2024). The six-item recognition test: mean 3.7 correct of 6, spam filtering identified correctly by 51 per cent, 11,004 US adults (Pew Research Center, 2023). Demographic differentials in face recognition: false positive rates varying by factors of 10 to beyond 100, 189 algorithms, 18.27 million images (NIST, NISTIR 8280). Credit score noise: model fit 16 per cent lower for the lowest income quartile, 50 million consumers (Blattner and Nelson, 2021). Lending disparities: 4.7 to 4.9 basis points on 2009 to 2015 purchase mortgages, 5.7 million loans (Bartlett et al., Journal of Financial Economics). The origin of Bayesian spam filtering as published machine learning, 1,789 messages (Sahami et al., AAAI 1998).
Strong, and it constrains what the level may claim. Far-right content was 0.17 per cent of YouTube viewing in 2016 and 0.30 per cent by 2019, with 55 per cent of arrivals at that content coming from off-platform links rather than recommendations, across 309,813 people and 21.4 million pageviews (Hosseinmardi et al., PNAS 2021). Switching the algorithmic feed to reverse chronological order for three months changed political attitudes by essentially nothing (US 2020 Facebook and Instagram Election Study).
Moderate. The Apple Card investigation: no evidence of unlawful discrimination found across roughly 400,000 New York applicants, with the regulator's own criticism that federal law requires lenders to explain denials and not credit limits (New York DFS, 2021). Wrongful arrests following a face recognition match: 14 named cases, most of them Black people, compiled by an advocacy organisation with no official national tally to check it against (ACLU, 2026).
Level 1, on generated answers
Strong. Legal fabrication: 58 per cent of queries for one model to 88 per cent for another, over 800,000 queries across 5,000 cases (Dahl et al., Journal of Legal Analysis 2024). The same corpus on calibration: all models systematically overestimate their confidence relative to their actual rate of fabrication. Human detection of machine-written text at 50 to 52 per cent across 4,600 participants and 7,600 texts, with neither payment nor feedback training helping (Jakesch, Hancock and Naaman, PNAS 2023). Synthetic faces identified at 48.2 per cent, below chance, and rated 7.7 per cent more trustworthy than real ones, three preregistered experiments (Nightingale and Farid, PNAS 2022).
Moderate. Verbalised confidence clustering between 80 and 100 per cent in multiples of five while measured accuracy in each band sits far below (Xiong et al., ICLR 2024). Fabricated citations at 55 per cent for one model and 18 per cent for its successor, 636 citations (Walters and Wilder, Scientific Reports 2023). Sourcing or accuracy problems in 45 per cent of news answers, 2,709 responses across 22 broadcasters and 14 languages (European Broadcasting Union and BBC, 2025). Attribution wrong on more than 60 per cent of 1,600 queries, with one tool wrong on 134 of 200 articles and hedging 15 times (Tow Center, 2025). Short-form factual questions: one model incorrect on 60.8 per cent and declining 1.0 per cent of the time, another incorrect on 19.6 per cent and declining 75 per cent, 4,326 questions (Wei et al., 2024).
Weak to moderate, and used for its gradient rather than its level. Citations entirely fabricated in 19.9 per cent of 176 cases with a mid-2025 model, rising from 6 per cent on the most common topic to 28 and 29 per cent on the two less common ones (Linardon et al., JMIR Mental Health 2025).
Moderate as argument, not measurement. Whether fabrication is reducible: two preprints reaching apparently opposite conclusions, neither peer-reviewed as of August 2026 (Kalai et al., 2025; Xu, Jain and Kankanhalli, 2024). Level 1 teaches the reconciliation rather than either position.
Level 2, on instructing the tool
Strong. Underspecified requirements are inferred correctly only part of the time and are roughly twice as likely to break on a model update, 240 prompts across three tasks (Yang et al., Findings of ACL 2026). Splitting one request across several conversational turns cost 39 per cent on average across every model tested, from analysis of more than 200,000 simulated conversations (Laban et al., ICLR 2026). Professional writing tasks: time down 0.8 standard deviations, graded quality up 0.4, and the correlation between a worker's first and second grade falling from 0.49 to 0.25, 453 professionals (Noy and Zhang, Science 2023). Customer support: 15 per cent more issues resolved per hour overall, about 30 per cent for the least experienced agents and no significant gain for the most skilled, 5,172 agents (Brynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025). Consultants: 12.2 per cent more tasks completed, 25.1 per cent faster, quality up 32 per cent inside the tool's competence, and 19 percentage points less likely to reach the correct answer on a task designed to mislead, 758 consultants (Dell'Acqua et al., Organization Science 2026). Kenyan entrepreneurs: no average effect, high performers about 15 per cent more profitable and low performers about 8 per cent less, 640 business owners (Otis et al., Management Science).
Moderate. Threats, tips and emotional appeals moved individual answers by up to 36 points in either direction and the aggregate by nothing, five models, 25 trials per question (Meincke et al., Wharton, 2025). Sixteen developers 19 per cent slower while believing themselves 20 per cent faster, with 9 per cent of their time spent reviewing generated output and fewer than 44 per cent of generations accepted (METR, 2025).
Weak, and named as such in the lesson. The "$335,000 prompt engineer salary" traced to the top of one range in one job advertisement. The tipping technique traced to a single social media post of December 2023 asserting it had been statistically checked, with no data ever published, and since falsified by measurement. The "115 per cent improvement from emotional prompts" traced to a relative gain on one benchmark in a 2023 technical report.
Level 3, on verification
Strong, and it is the argument for the whole level. Help from a language model raised fact-checkers' accuracy to 74 per cent on average and dropped it to 35 per cent when the model's explanation was wrong, below the 49 per cent achieved unaided, 80 participants across five conditions and 1,500 annotations (Si et al., NAACL 2024). Radiologists' accuracy fell from 79.7 per cent to 19.8 per cent for the least experienced and from 82.3 per cent to 45.5 per cent for the most experienced when the automated suggestion was wrong, 27 readers, 50 cases, randomised presentation order (Dratsch et al., Radiology 2023). Participants agreed with incorrect predictions about 80 per cent of the time in a task where checking was possible by hand, 731 participants across five experiments (Vasconcelos et al., CSCW 2023). Inoculation retention of 100 per cent with repeated testing and 36 per cent without (Maertens et al., 2021). Media literacy transfer of 26.5 per cent in the United States, 17.5 per cent among educated urban Indians and nothing in rural India (Guess et al., PNAS 2020). No average effect from an hour of in-person media literacy training in India, and a backfire among ruling-party supporters, 1,224 people (Badrinathan, American Political Science Review 2021).
Strong as corrections of circulating figures. The 90th-percentile bar examination claim corrected on peer review to below the 69th percentile against a July cohort and roughly the 15th percentile on the essays (Martínez, Artificial Intelligence and Law 2025). The prediction that 95 per cent of customer interactions would be powered by artificial intelligence by 2025, traced to a vendor with no published methodology, against 33 per cent of European adults having used a generative tool at all.
Weak, and used as the worked example of an unsourced number. The claim that fabrication cost businesses $67.4 billion in 2024, traced through four commercial pages citing each other in a circle and terminating at an affiliate marketing site that could not be opened.
Level 4, on privacy and impersonation
Strong. Pooled human deepfake detection at 55.54 per cent, and audio specifically at 62.08 per cent with a confidence interval from 38 to 83 per cent, across 56 papers and 86,155 participants (Diel et al., 2024). Primed and familiarised listeners still missed 27 per cent of speech deepfakes, and accuracy rose to 85.59 per cent only when a real and fake clip were compared side by side, 529 participants (Mai et al., PLOS ONE 2023). Cloned voices identified as artificial 60.8 per cent of the time, and accepted as the same speaker as the real person at a median of 83.3 per cent, 604 listeners (Barrington, Cooper and Farid, Scientific Reports 2025). Provider data policies, quoted verbatim and dated, from the providers' own pages. The preservation order in the consolidated copyright litigation, read from the order itself (Wang, S.D.N.Y., 13 May 2025).
Moderate. The Hong Kong engineering-firm fraud: the company confirmed that fake voices and images were used and never confirmed the amount, which rests on a police press briefing. Nigerian deepfake incidents: the securities regulator's September 2025 warning about fabricated endorsements, the advertising regulator's confirmation of an artificial video of the President promoting a scheme, and the World Trade Organization Director-General's own statement that a viral investment video of her was manipulated. All are official or first-person, and none carries a victim count or a loss figure.
Weak, and reported as self-perception rather than incidence. One in four adults reporting direct or second-hand experience of a voice scam, 7,054 respondents in seven countries, none of them African, sponsored by a company selling protection against the threat measured (McAfee, 2023).
Not used. The 2019 European energy-company voice fraud, which has one source, the victim's insurer, speaking to a newspaper: no company named, no suspect, no court record in seven years. Level 4 uses it as an example of citation hygiene rather than as a case.
Level 5, on judgement
Strong. Pooled human and artificial intelligence combinations performing worse than the better of the two alone at Hedges' g = -0.23, with decision tasks at -0.27 and creation tasks at +0.19, across 106 studies and 370 effect sizes (Vaccaro, Almaatouq and Malone, Nature Human Behaviour 2024). Individual novelty up 8.1 per cent with five generated ideas while the resulting body of stories became measurably more similar to each other, 293 writers and 600 evaluators (Doshi and Hauser, Science Advances 2024). Users' calibration about a model worse than the model's own calibration about itself, and longer explanations raising user confidence without raising accuracy, six experiments (Steyvers et al., Nature Machine Intelligence 2025).
Moderate. Extensive editing of generated drafts, an average of 300 edit actions per abstract, leaving the distributional differences between generated and human text intact (Queiroz Da Silva et al., 2026).
Weak, and named as such. Two hours of rework per instance of unfinished-looking generated work, from a self-report survey of 1,004 desk workers published as a magazine article by a company selling coaching, whose dollar figures are extrapolations from recalled minutes.
Claims removed from this course, and why
"Pew found that most people interact with artificial intelligence constantly without realising it." No such finding exists. Pew's figures run the other way: 27 per cent of American adults in 2022 and 2024, and 36 per cent in 2026, said they interact with it several times a day. The recognition-gap claim the course needed is real and belongs to Gallup. The composite was removed and both sources are now cited separately. It is used in Level 3 as a worked example of a plausible, correctly-attributed-looking statistic that decays into error through retelling.
"88 per cent of Nigerian adults use artificial intelligence chatbots." Circulating in technology press attributed to a vendor study. The primary document could not be made to yield that figure. Removed.
"One in five Nigerians lost money or data to scams." Available only in secondary reporting of a report behind a registration wall. Removed pending a primary source.
"Only 0.1 per cent of people can detect deepfakes." A vendor's perfect-score rate across a mixed stimulus set, circulated as a detection rate. Replaced with the pooled meta-analytic figure of 55.54 per cent.
"A $21 million grandparent scam used artificial intelligence voice cloning." Three United States Department of Justice releases and one immigration enforcement release on this prosecution were read in full. None mentions artificial intelligence, voice cloning or synthetic voice. The attribution appears only in secondary reporting and in an incident database title. It is used in Level 4 as the cleanest available demonstration of how an artificial intelligence attribution attaches itself to a case where the government never alleged one.
"Deloitte's forecast of $40 billion in artificial intelligence-enabled fraud by 2027." A consultancy projection built partly from its own practitioners' judgement, routinely quoted as a measured loss. Not used as a figure.
"The 23 minutes and 15 seconds it takes to refocus after an interruption." Not in this course's subject matter, and included here because it is the same failure: a precise figure with no paper behind it. Level 3 teaches the trace rather than the number.
Worked exemplars
Four artefacts, annotated. Each one is what the output of an exercise in this course looks like when it meets the standard the gates ask for.
Reading a completed artefact is not the same as producing one, and this page is the smaller half of the work. It is here because three of the four exercises in this course produce something many readers have not seen: a kept record of a tool being wrong.
The annotations say why each choice was made and, more usefully, what the weaker version of the same artefact looks like. The weaker version is nearly always shorter, tidier, and missing the part that would have been embarrassing.
What these are, and what they are not
These four were written to the standard rather than collected from participants. No learner produced them. The arithmetic ties, the dates are consistent, and the failures they record are the failures the evidence in this course predicts.
That is a provenance state, not an educational defect. A constructed artefact is a research-informed worked example: it makes expert reasoning, evidence handling and the shape of a good record visible. Real learner artefacts can later enrich and validate the set, with permission and names changed, while these examples remain legitimate teaching assets marked as constructed.
What is open, and what is not
The free levels of this programme are readable with no account at all — the real levels, not samples. An account carries your progress, your gate claims and your saved work. The remaining levels, the tools, the gates and this programme’s worked exemplars are opened together when you enrol.
Create an account