Every Agent Demo Works. Then You Give It Real Work.
Aug 6, 2026 · 5 min read · by Jordan Kwan
TL;DR: Agent demos and agent products are different artifacts. The demo is sampled from the happy path; production is sampled from reality, where per-step errors compound until a 95%-reliable step becomes a 36%-reliable twenty-step job. I checked 19 agent product homepages: 10 lead with a demo, 6 publish any accuracy number, and exactly zero publish an error rate. Judge the failure behavior, not the happy path.
You have seen the demo. Everyone has seen the demo. The agent reads the email, checks the calendar, books the meeting, updates the CRM, and the room exhales. Then you buy it, point it at your actual inbox, and week two produces a confidently wrong reply to your biggest customer that nobody catches for four days.
This is the most cross-validated pattern in AI right now, and I do not think it is a bug that the next model release fixes. The gap between demo and deployment is structural.
What is the pattern?
On August 5, 2026 I went through 19 AI agent product homepages, one by one: 11x, Artisan, Lindy, Devin, Sierra, Decagon, Harvey, Fin, Ada, Zapier Agents, and the rest of the current crop. 10 of 19 lead with a demo. Video, audio, animated walkthrough, an agent booking a meeting in a loop, forever. 6 of 19 publish any quantitative accuracy or resolution figure at all, and those are cherry-picked customer numbers, resting on a definition of "resolution" that counts a customer who gave up as a win. Zero of 19 publish an error rate. Not one vendor selling an autonomous worker will tell you, on the page where they sell it, how often it is wrong.
The one error-adjacent number I found anywhere was a comparative hallucination-reduction claim. One vendor even cites "87% of AI projects never reach production" as a reason to buy theirs, a cousin of the 95, the 80 and the 40 that get quoted interchangeably and measure three different things. The industry's marketing knows exactly what the industry's product statistics are. When a failure record does get written, it is usually by outsiders: Comet's six security disclosures are the best-documented example.
Why does the demo always work?
Because the demo is not a sample of the product. It is a sample of the five cases the product was rehearsed on. Production is a sample of your actual distribution, which contains case six, and case six is where an agent improvises. When I called agents cron jobs with anxiety, this was the mechanism: a loop, some tools, and no confident stopping rule.
The math is the unforgiving part. Utkarsh Kanwat, who builds production agent systems for a living, laid it out plainly: at 95% per-step reliability, optimistic for current models, a 5-step workflow succeeds 77% of the time, a 10-step workflow 59%, and a 20-step workflow 36%. The demo is 3 steps. Your job is 20.
Benchmarks that force repetition agree. Sierra's own tau-bench found a top agent scoring above 60% on a retail task once, but below 25% when asked to succeed 8 times in a row on the same task. And when Mercor's APEX-Agents benchmark put frontier models on long-horizon professional tasks at its January 2026 launch, the best scored 18-24%, with models failing every single criterion on 40-62% of tasks. The August 2026 leaderboard has climbed to roughly 40-43%, which is real progress and still a coin flip on work an employer would fire a human for coin-flipping.
So why do teams keep buying the demo?
Because the demo is the only evidence on offer, and the failure data lives elsewhere. It lives in Gartner's April 2026 survey of 782 infrastructure and operations leaders, where 20% of AI projects fail outright and only 28% fully meet ROI expectations. Gartner research director Melanie Freeze named the cause: "The 20 percent failure rate is largely driven by AI initiatives that are either overly ambitious or poorly scoped." It lives in Teradata's 2026 "Arrested Automation" study, where 78% of enterprises have an agent pilot running and 14% have scaled one to organization-wide use. And it lives on Reddit, where an operator who talked to companies deploying agents wrote the whole post in one line: "90% of agents break after launch and no one talks about that." His follow-up is better: "You basically need to babysit these agents like they're interns who lie on their resumes." And the thesis of this post, from the same thread: "Builders love Demos, buyers don't care."
Andrej Karpathy, asked on the Dwarkesh Podcast why you cannot simply hire an agent like an employee, gave the least-hedged answer in the industry: "The reason you don't do it today is because they just don't work." His timeline for working through the deficits is a decade, not a release cycle. (OpenAI now sells exactly that employee; the early ChatGPT Work record is the live test of his claim.)
None of this means agents are useless. The scoped slice genuinely works: one job, tight tools, a human on the trigger. That is how I run the drafting agent in my own Reachium stack, and it is the configuration every surviving deployment in the data seems to converge on. What does not work is the thing the demo sells, which is the unsupervised twenty-step employee.
What would change my mind?
Three things, any of which I would genuinely celebrate. A vendor publishing its error rate on the homepage, next to the demo, would end this take on the spot; the first one to do it will also corner the enterprise market, because buyers are not stupid twice. A benchmark like APEX crossing 80% pass@1 on long-horizon professional tasks would mean the compounding problem is actually being solved rather than survived. And a Teradata-style study showing most pilots scaling to production would mean the demos started sampling reality.
Until one of those happens, the operator move is simple: when the demo ends, ask the only question that matters. "Show me it failing, and show me what happens next." The vendors with a real product have an answer. The vendors with a demo have another demo.
Written by Jordan Kwan, founder of Reachium.
I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.
See what Reachium does ↗