What the capability numbers actually say
Two headlines circulate, and both are wrong. One says agents can do most office work now. The other says agents cannot do real work at all. Holding either belief will cost you: the first through failures you did not supervise for, the second through capacity you never built. The evidence supports neither, and the spread between the benchmarks is itself the lesson.
Start with the harshest number. TheAgentCompany benchmark placed agents inside a simulated software company and gave them 175 tasks across project management, HR, finance, administration and engineering, in an environment with colleagues to message, websites to navigate and files to find. The best 2025-generation model completed 30.3 per cent of tasks fully, 39.3 per cent with partial credit (Xu et al., 2025). The worst categories were data science, administrative and finance, where several models scored zero. Three failure patterns recurred: missing social context, such as not grasping what a colleague's message implied; failing at web interfaces, clicking and navigating badly; and self-deception, fabricating shortcuts or claiming completion when stuck. The figures are 2025-generation and should be treated as a floor that has since risen. The failure patterns are the durable finding: they have survived model generations.
Now the friendliest number. WorkBench tests structured tool-calling over clean workplace databases: email, calendar, analytics, CRM, 690 tasks. The 2024 GPT-4 baseline completed 43 per cent, with harmful side effects, wrong deletions and wrong sends, in 26 per cent of attempts. The best 2026 model completes 88.8 per cent with 2.5 per cent harmful side effects (Styles, 2026). That is a real and steep improvement. Even at the top, two errors persist: misinterpreting retrieved data, and treating truncated search results as complete, acting on the first page as if it were the whole.
Between them sits Gaia2: 1,120 scenarios in an event-driven simulated phone environment where messages arrive and conditions change mid-task (ICLR 2026). The best frontier model passes 42.1 per cent on the first attempt, and scores 0 per cent in default mode on time-sensitive tasks, because reasoning depth trades directly against latency: the model thinks well or answers on time, not both (Froger et al., 2026).
So the same generation of systems scores in the high eighties on one benchmark and a third or less on two others. The variable is not difficulty. WorkBench tasks are not easy; some involve multi-step operations across several tools. The variable is structure. Clean inputs, defined tools, checkable outputs: high eighties. Ambiguity, visual interfaces, social context, timing pressure: a third or less. Capability is jagged, and it is jagged along the axis of structure.
What you do differently: classify your own tasks by structure, not by difficulty. A task you find intellectually hard, comparing supplier quotes against written criteria, may be highly structured and sit in delegable territory. A task you find trivial, replying to an ambiguous two-line email from a long-standing client, is saturated with social context and sits in the weakest zone. Your sense of difficulty was calibrated on human cognition and will mislead you here. For each task in this week's inventory, ask: are the inputs clean or ambiguous, are the tools defined or improvised, is the output checkable or a matter of feel, does timing matter. Four structure questions predict delegation success better than any impression of hardness.
Do not build plans on the specific figures in this lesson, and do not dismiss the lesson when newer figures appear. Every score here will be stale within a year. The failure patterns, missing social context, interface fragility, truncation blindness, claiming completion when stuck, have persisted across generations, and your supervision design in Levels 4 and 5 targets patterns, not scores. Plan for the patterns; let the numbers move.
The same generation of systems scores 88.8 per cent on WorkBench and roughly 30 per cent on TheAgentCompany. What best explains the spread?
TheAgentCompany's 30.3 per cent is a 2025-generation figure. How should you use it when planning in 2026?
Notes are kept with your account, alongside your progress and your gate claims. The lesson itself is readable without one.