Agents’ Last Exam Shows the Reliability Gap Behind Today’s AI Agents
Agents’ Last Exam moves AI evaluation from short benchmark tasks toward real professional workflows — and finds that reliable, end-to-end completion remains the hard part.

AI agents can look remarkably capable when a benchmark asks them to write a function, answer a question, call a tool, or complete a short sequence of actions. Agents’ Last Exam asks a harder question: what happens when the unit of evaluation starts to resemble an actual piece of professional work?
The answer is much less flattering. The benchmark, developed by UC Berkeley RDI with hundreds of domain experts, evaluates agents on long-horizon, economically valuable workflows with outputs that can be checked against explicit success criteria. In the paper, the hardest tier has an average full-pass rate of just 2.6% across mainstream model-and-harness configurations. That does not mean AI can perform only 2.6% of jobs. It means that when the benchmark demands a complete, correct deliverable across difficult multi-step work, frontier agents still fail far more often than headline benchmark scores might suggest.
The benchmark is trying to measure work, not intelligence in isolation
Most AI benchmarks deliberately simplify reality. That is useful when researchers want to isolate a capability such as mathematics, coding, retrieval, or tool use. It is less useful when a company wants to know whether an agent can take ownership of a workflow and return something that is actually ready to use.
Agents’ Last Exam, or ALE, is designed around that gap. The paper organizes more than 1,000 tasks across 55 subfields and 13 industry clusters, grounded in the U.S. O*NET/SOC occupational taxonomy. The project has continued to grow since the paper was posted: its public site now describes more than 1,500 collected tasks and participation from more than 300 industry experts.
The environment matters as much as the task list. ALE places agents inside Windows or Linux sandboxes with the professional software, files, web access, terminal tools, and graphical interfaces required by the workflow. Tasks use real input data and hidden references, and the resulting artifacts are graded against verifiable outcomes. An agent is not rewarded simply for producing a plausible explanation of what it would do. It has to leave behind the right result.
Why full-pass changes the picture
That distinction makes full-pass rate a revealing metric. In a long workflow, partial competence can still produce an unusable deliverable. A financial model with one broken formula, a design file missing a required object, a medical analysis built from the wrong data field, or a report that omits a required section may contain substantial correct work while still failing the job.
The ALE team highlights a recurring failure mode: agents often announce that a task is finished before they have actually verified the output. The work may be missing files, contain the wrong counts, omit required fields, or violate constraints even though the agent ends with confident language saying the checks passed. This is less a problem of generating a smart next step than of maintaining discipline across an entire trajectory.
That helps explain the 2.6% figure. The hardest tier is intentionally difficult and should not be treated as a representative sample of everyday office work. But it reveals a compounding effect that short benchmarks often hide: planning errors, tool mistakes, misunderstood requirements, context drift, and weak verification become more damaging as the number of steps increases. A model can be excellent at most individual actions and still be unreliable at the workflow level.
The agent harness helps — but it does not erase the model gap
One of ALE’s useful follow-up experiments separates the model from the software harness around it. A harness manages the agent loop: it builds context, exposes tools, records observations, handles memory, and decides how the model continues working. In principle, a sophisticated harness could compensate for weaknesses in the underlying model.
The project’s June analysis found that model choice moved ALE full-pass performance much more than harness choice. With OpenClaw held fixed, changing the model produced an 18-point pass-rate spread. Holding a model fixed while changing harnesses produced spreads of roughly five to six points. That does not make orchestration irrelevant, but it suggests that today’s reliability ceiling is still driven heavily by the underlying model rather than by adding more layers around it.
The cost result is just as interesting. Berkeley RDI built a leaner harness called ALE-Claw and compared it with OpenClaw using the same GPT-5.5 model. ALE-Claw reached the same accuracy band while using 44% fewer input tokens, costing 41% less, and taking 60% less wall-clock time. Extra agent machinery can consume context and money without necessarily changing whether the final artifact passes.
This is why “cost per completed task” is becoming the useful metric
That finding connects agent evaluation directly to economics. A provider can advertise a low token price, a large context window, or a strong benchmark score, but an enterprise ultimately pays for outcomes. A failed run still consumes tokens, tool calls, compute time, and employee attention. A nearly correct deliverable may require enough human repair that automation saves little.
A better production metric is therefore total cost per successful, reviewable task. That includes inference, retries, runtime, external tools, verification, and human intervention. ALE does not solve that measurement problem by itself, but it moves evaluation closer to it because success is attached to a complete artifact rather than a model’s answer to an isolated prompt.
The benchmark has important limits
ALE should not be read as a direct forecast of employment or automation. It focuses on non-physical work that can be executed in computer environments, and its need for verifiable grading naturally favors workflows whose outputs can be checked objectively. Many real jobs depend on negotiation, accountability, tacit knowledge, interpersonal judgment, physical action, or ambiguous success criteria that are difficult to represent in a sandbox.
It is also a living benchmark. Tasks are added, models change, harnesses improve, and the public leaderboard moves. The paper’s 2.6% hardest-tier average is a snapshot tied to a particular set of configurations and should remain labeled as such. Newer systems can improve rapidly without invalidating the central observation: long-horizon professional work exposes failure modes that short tests systematically underweight.
What to watch next
The most important signal will not be whether ALE’s leaderboard rises. It almost certainly will. The better question is how it rises. Do agents become better because the models reason and verify more reliably, or because increasingly specialized scaffolds teach them how to survive the test? Do higher pass rates arrive with lower cost, fewer retries, and less human supervision? And do gains transfer across finance, engineering, medicine, design, research, and other domains rather than concentrating in the kinds of tasks AI already handles well?
Agents’ Last Exam is valuable because it changes the unit of measurement. The industry has spent years asking whether a model can answer correctly. The commercial question is whether an agent can take responsibility for a piece of work, operate the necessary tools, catch its own mistakes, and deliver something another person can trust. On that exam, frontier agents are improving quickly — but they have not graduated yet.
Sources and further reading
- Agents’ Last Exam - arXiv
- Agents’ Last Exam — Project Site - UC Berkeley RDI
- Agents’ Last Exam Evaluation Framework - UC Berkeley RDI / GitHub
- Does the Harness Matter? Lessons from ALE-Claw on Agents’ Last Exam - UC Berkeley RDI