The Agent Reliability Numbers Nobody Wants to Put in a Pitch Deck

A cluster of builders spent the weekend circulating the same uncomfortable statistics: most agent pilots die before production, and the best agents still fail the majority of real office tasks.

For all the swarm enthusiasm, a quieter and more honest conversation ran in parallel this weekend — one about how badly agents still perform when the demo is over. The headline number came from @theblankcollar, citing a Carnegie Mellon benchmark in which "the best AI agent failed ~70% of real office tasks." Not edge cases. Real office work.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.