The AI Practicum
Learn to work with AI for real tasks, better decisions and better results.
- Levels
- 6
- Lessons
- 30
- Knowledge checks
- 60
- Gates
- 61
- Working tools
- 7
- Simulations
- 2
- Worked exemplars
- 5
◆ Fourteen sessions over eight weeks
If you have completed Making Sense of AI, this is your practical next step: from first familiarity to using AI effectively and efficiently on real work. If AI still feels unfamiliar or uncertain, you are welcome here too. You will learn to choose tasks that suit AI, give clear direction and requirements, produce and refine useful work, check important claims, protect confidential information, and remain responsible for what you send out. You will also measure whether AI actually saved time or improved quality.
The course follows one real task that you do repeatedly and that produces something another person reads or acts on. You will name that task, work on it throughout the course, and ask a named reader or judge to assess the results. That through-line keeps the practice concrete: you can see what changed, what helped, and what still needs your judgement.
Three capabilities, built in order.
Scoping means deciding what the work really is, which parts are a good fit for AI, and which parts need your own knowledge or another source. Direction means describing the outcome, materials, constraints and checks clearly enough for a model to help in useful stages. Custody means staying responsible for the result: reviewing it, verifying what matters, protecting information, and deciding what you are willing to sign.
By the end, you will have a measured baseline and re-measurement of the same kind of work, a frontier/task-fit map showing where AI helped and where it did not, five reusable specifications, review and verification records, and one real piece of work completed end to end. These are working tools and evidence you can carry into your next task, not a collection of magic prompts.
Research supports the course without taking it over. Studies of human-AI work, requirement-setting, review and verification help shape the exercises and the questions you ask of your own results. The detailed evidence and its limits are available in the lessons and Sources section; here, the focus stays on making better work with care.
The levels
- 0Measure what you actually getTake the two numbers this course is judged against, and find out that you cannot read either of them off your own experience.Free
- 1Define the work, and find your frontierSpecify what you want well enough that somebody else could check it, then test which parts of your own work sit inside the model's competence rather than guessing.Locked
- 2Brief, and what briefing actually buysStop doing the things that have been tested and do not work, and understand why the things that failed felt like they were working.Locked
- 3Review what came backLook at the output as somebody who expects to find something wrong in it, using a procedure rather than an impression, because the impression is the thing that has been shown to move independently of the answer.Locked
- 4VerifyCheck the claims that carry weight against something outside the model, using a procedure built for the error class that exists now rather than the one that existed three years ago.
Your working ground
◇ Before you start
Four things get fixed before Level 0 and do not change for eight weeks. Everything you measure is measured against them.
One: the recurring task.
Name one professional task you perform repeatedly, at least weekly, that produces something another person reads or acts on. A weekly report. Meeting minutes. Client correspondence. A funding note. A lesson plan. Grant reporting. A shift handover.
It must be recurring, because the second measurement in week eight has to be comparable to the first. It must produce an artefact, because "thinking about strategy" cannot be graded. And it must have a reader, because your own opinion of your output is the thing this course is built to distrust.
Write the task, the frequency and the reader's name.
Two: the judge.
One person, named, who will look at two pieces of your work eight weeks apart and answer four questions about each. They do not need to know which was which, and it is better if they do not.
They must be somebody who normally receives or reviews this kind of work, because a stranger cannot tell whether a report is right, only whether it reads well, and Level 3 is entirely about the difference.
Three: the ground-truth task.
One task you do where you can find out afterwards whether you were right. A forecast that resolves. A figure that gets audited. A recommendation whose outcome you observe. A diagnosis that gets confirmed.
This is the hardest of the four to name and the most important. In a re-analysis of the 74-study human-AI corpus, only 14 per cent of experiments gave participants outcome feedback, and studies with feedback showed synergy at g = 0.12 against g = −0.17 without it (Berger et al., 2026). Most professional work has no feedback loop at all. You write the memo and never learn whether the advice was right.
If you genuinely have no such task, write that down. It is a finding about your working life and it changes what Level 5 asks of you.
Four: the confidential boundary.
Before week one, write down what you are not permitted to put into a tool your organisation does not control. Client identities. Personal data. Unpublished figures. Anything under an agreement.
Get this from the actual policy if one exists. In a survey of 48,340 people across 47 countries, only about two in five reported that their organisation had a policy at all, almost half reported uploading sensitive information, and over half reported concealing their AI use (Gillespie et al., 2025). If yours has no policy, write your own boundary and note that you wrote it.
Rules of practice
§ Non-negotiable
Seven rules. Four have evidence behind them and are marked. Three are design choices for this course and are marked as such, because a course about not accepting unsupported claims should not open with unsupported claims.
One. Every exercise runs on your own named recurring task. Nothing in this course is practised on an invented scenario. Design choice. The reason is that the jagged frontier is task-specific and unpredictable, so a generic exercise cannot tell you where yours runs.
Two. Output is judged by somebody other than you. Evidenced. Across four separate measurements, developers' predictions and retrospections about their own AI-assisted speed were wrong by roughly 39 percentage points in the same direction, immediately after the work (METR, 2025). Self-assessment is not available as an instrument here.
Three. You record what you kept. For every AI-assisted output, record the proportion you retained unchanged. Evidenced. In the professional-writing experiment, 68 per cent of treated participants submitted the model's output without editing it, average post-paste working time was 3 minutes, and the correlation between post-paste time and final grade was approximately r = 0.00 (Noy & Zhang, 2023). In the consulting experiment, the modal retention score was about 0.87, and prompt training increased copy-pasting rather than reducing it.
Four. Important claims are verified against the source, not against the model. Evidenced. Asking a model to check itself does not work: across ten models and 67,640 trials, following up with a challenge like "are you sure?" reduced accuracy by an average of 17 per cent (Laban et al., 2025). Intrinsic self-correction without external feedback degrades reasoning performance (Huang et al., ICLR 2024).
Five. Retrieval does not count as verification. Evidenced. Purpose-built legal research tools with retrieval over authoritative databases still hallucinated on 17 per cent of queries (Lexis+ AI) and 33 per cent (Westlaw AI-Assisted Research) across 202 preregistered queries (Magesh et al., Journal of Empirical Legal Studies, 2025). Grounding reduces the rate. It does not remove it.
Six. You keep the source reachable. A summary does not replace the document it came from, and any claim that matters has to be traceable to something you could open. Design choice, and the reason is in Level 4: the failure mode in summarisation is omission rather than fabrication, and omission is invisible in the output.
Seven. You are the author. Nothing goes out that you would not sign. Design choice, and it is also the position taken by every professional body examined for this course. The American Bar Association's Formal Opinion 512 puts it as a duty: relying on a generative tool "without an appropriate degree of independent verification or review of its output" can violate the duty of competence.
Assessment
◈ Reference
Six components. The weights are a design choice; what each component demands is not negotiable.
Baseline and re-measurement, 15 per cent. Two pieces of work on the same recurring task, week one and week eight, with the judge's four answers on each and your retention proportion recorded for both. Marked on whether the measurement was actually taken and taken the same way twice, not on whether the numbers improved.
Frontier map, 15 per cent. A written record of at least eight subtasks within your recurring work, each placed inside or outside the frontier on evidence rather than expectation, with at least two that moved after you tested them.
Specifications, 20 per cent. Five reusable task specifications, each used at least twice on real work, each stating the objective, the material, the constraints, the output shape and the check. Marked on whether a colleague could run them without asking you a question.
Review record, 20 per cent. For at least six AI-assisted outputs: what you changed, what you kept, and one error you caught. Including at least one output you discarded entirely, and at least one where the fluent version was the wrong one.
Verification log, 15 per cent. At least ten claims checked against a primary source, with the time each check took recorded. Including at least one claim that did not survive.
Final project, 15 per cent. One real task end to end, with the specification, the stages, the first output, the corrections, the verified claims, the disclosure decision, and the final approved version. Plus what the model did well, what it did badly, and what remained a human decision.
Nothing is assessed on the elegance of a prompt.
Sources
† Evidence
Every figure in this course carries a citation and an access status. Three grades are used in the text.
Evidenced means a meta-analysis or several independent studies, or one large study whose full text was read.
One study means exactly that, with the sample size stated so you can weigh it.
Convention means no test exists. This course contains conventions, because some questions have not been studied and the honest response is to say so at the point of use rather than to teach silence.
A structural warning about this evidence base. Two systematic surveys of prompting techniques catalogue 58 and 29 techniques respectively. The one adequately powered independent replication attempt covered six of them. The ratio of proposed techniques to independently replicated techniques in this field is roughly 58 to 2. Treat any confident claim about how to talk to a model against that background. This course was built against it: no phrasing technique is taught here on the strength of an unreplicated claim, and where the course takes a position of its own, it says so at the point of use and gives the mechanism, so you can test it on your own work rather than take it on trust.
A second warning, about staleness. Every measurement here has a model version attached where the finding depends on one. Hallucination rates in particular have moved substantially between model generations, and a verification protocol built for 2023 output does not catch 2026 output: frontier research agents now resolve 94 to 100 per cent of the links they cite while only 39 to 77 per cent of citations actually support the claim they are attached to, and factual support falls by roughly 42 per cent as tool use rises while the surface indicators stay clean (Onweller et al., 2026).
Claims removed from this course, and why.
"AI makes everyone more productive." Removed. It depends entirely on whether the task sits inside the frontier, and on real tasks with expert practitioners it has been measured going the other way at 19 per cent slower.
"Prompt engineering is a core skill." Removed. The one controlled comparison of curricula found requirement-specification training producing 20 per cent gains against 1 per cent for technique training, and the technique nulls are extensive.
"Chain of thought improves reasoning." Qualified rather than removed. It improves maths and symbolic reasoning substantially and everything else by 0.7 points, and on reasoning-trained models it is roughly neutral or negative.
"Assign the model a role or persona." Removed. 162 personas, 2,410 questions, nine models, no significant improvement.
"Tell it the task is important to you, or offer it a tip." Removed, and it is worse than null. In one study across five models with 4,950 runs per condition, "this is important to my career" produced −4.0 percentage points, p = 0.002.
"Ask the model to check its own work." Removed. See Rule Four.
"Your brain on ChatGPT." Removed. The widely circulated EEG study rests on 54 enrolled participants with 18 in the crossover the conclusion depends on, up to a thousand analyses of variance, and a control condition that contradicts its own proposed mechanism. Level 5 uses the three controlled experiments that came later instead.
The Centaur and Cyborg working styles. Removed as a recommendation. They are descriptive patterns observed in interaction logs. No trial has trained anyone into either and measured the result.
Worked exemplars
‡ Worked to standard
Five real published artefacts, reproduced with their operative wording and their limits attached. They are not illustrations. Three of them impose duties that already apply to you.
The competence opinion. American Bar Association Formal Opinion 512, issued 29 July 2024, on generative AI. It is reproduced because it states in one document what a professional body considers the minimum: independent verification, the confidentiality problem with tools that train on input, when you have to tell the person you are working for, and what you may charge for.
The certification order. A United States federal judge's standing order requiring any party to certify either that no part of a filing was drafted by generative AI or that any language so drafted was checked by a human. The rationale paragraph is reproduced verbatim because it is the clearest short statement of why the fluency is the problem.
The statute. Articles 4, 5, 14 and 50 of the EU AI Act: AI literacy, prohibited practices, human oversight, and transparency obligations. Reproduced with the current implementation timing, including which parts were softened in 2026.
The framework. The NIST AI Risk Management Framework and its Generative AI Profile, with the definitions of confabulation and human-AI configuration quoted, because those two definitions do more work than most training material.
The authorship rule. The ICMJE recommendation and a major publisher's policy on AI in manuscripts. Reproduced because both resolve the same question the same way: the tool cannot be an author, use must be disclosed, and the human is responsible for accuracy. That is the position this course takes on every piece of work you produce.
What is open, and what is not
The free levels of this programme are readable with no account at all — the real levels, not samples. An account carries your progress, your gate claims and your saved work. The remaining levels, the tools, the gates and this programme’s worked exemplars are opened together when you enrol.
Create an account