What happens when an AI that is 95% right is given permission to act 100 times?

Treat the numbers as illustration, not a lab measurement. A model that is usually right in chat still faces a different question once it may call tools, edit files, spend budget, or change a customer’s world. Capability is how impressive the reasoning looks. Reliability is whether the job still finishes safely when the steps grow in a long-running workflow (here, 100 steps).

From Models to Agents Part 5 named the uncomfortable compound: wrong observations become wrong plans. AI Right Now Part 8 covered generative failure modes. This series is the reliability path — failure modes, benchmarks that measure different things, compounding math, human over-trust, and a sharper definition of useful intelligence: knowing when not to act.

OSWorld 2.0 makes the gap concrete for computer-use agents: workflows that take a human on the order of hours still expose hidden state, mid-task change, constraint drift, and skipped verification — even when the same systems look strong on shorter demos.

What you will get

Five parts. One thesis: the next phase of agents may be about making them boringly reliable.

The series

  1. The illusion of intelligence — why reasoning demos create unrealistic expectations
  2. The seven ways agents fail — a non-exhaustive trajectory checklist, plus why small per-step error becomes huge end-to-end risk
  3. The benchmark illusion — SWE-bench, OSWorld, METR, ReliabilityBench, and multilingual gaps
  4. The human failure mode — automation bias, over-trust, ambiguous responsibility
  5. Knowing when — and what reliable agents need — when to act, verify, stop, and ask; the reliability stack; the closing distinction

Sources