Hype Index
Signal

Hype Index: 'AI Agents Can Now Work 14 Hours Autonomously'

Aug 15, 2026 · 6 min read · by Jordan Kwan

TL;DR: METR's time horizon is a 50% success rate: the task length at which a model succeeds half the time. On METR's own published data file, Claude Opus 4.6 sits at 718.8 minutes (11h 59m) at the 50% threshold and 69.9 minutes (1h 10m) at the 80% threshold, a gap of 10.3x. I fetched 27 articles citing this work and 25 loaded; 8 of them never mention the 50% threshold, and 17 never name the task suite. The chart is real. "14 hours" is not METR's number for anything, and there is no single doubling time: METR publishes three.

What is the claim?

Stated the way it actually circulates, stripped of the chart:

AI agents can now work autonomously for 14 hours. The doubling time is 105 days. That is a 240x jump.

All three of those numbers are downstream of one source, METR's Time Horizon 1.1, published 2026-01-29. None of the three appears in it.

What does METR actually measure?

A time horizon is a length of human expert work paired with a success rate. METR's live dashboard, last updated 2026-05-08, defines it in one sentence: "the 50%-time horizon is the duration at which an agent is predicted to succeed half the time."

That threshold is the entire ballgame, and you can see why in METR's own raw numbers. The dashboard links a machine-readable file, benchmark_results_1_1.yaml, with 26 models in it. I parsed it. The values are in minutes, which I confirmed against the Time Horizon 1.1 appendix, where the identical Opus 4.5 cell is written out as "289 mins [104, 1285]".

Model (release date) 50% horizon 80% horizon
Claude Mythos Preview, early (2026-04-07) 1,044.8 min (17h 25m) 185.9 min (3h 06m)
Claude Opus 4.6 (2026-02-05) 718.8 min (11h 59m) 69.9 min (1h 10m)
Claude Opus 4.5 (2025-11-24) 293.0 min (4h 53m) 49.4 min
GPT-4o (2024-05-13) 7.0 min 1.3 min

Opus 4.6 at 80% reliability is a 70-minute model. At 50% it is a 12-hour model. Same model, same evaluation, same page, and a factor of 10.3 between them depending on which line you read. Every horizon figure is meaningless without the threshold attached, which is why METR always attaches it and almost nobody downstream does.

What did the count show?

I collected 27 named 2026 pieces citing METR's horizon numbers: news write-ups, vendor blogs, newsletters, forecaster posts and research explainers. Two returned HTTP 429 and never loaded (aiwiki.ai and lesswrong.com), so the denominator is 25, not 27. I did not substitute replacements to keep a round number.

Of the 25 that resolved:

  • 17 state the 50% success definition. 8 do not. Those 8 print an hours figure with no threshold anywhere on the page.
  • 8 name the task suite (HCAST, RE-Bench or SWAA). 17 never say what the model was tested on.
  • 13 mention the 80% horizon. 12 never acknowledge a second threshold exists.
  • 17 link metr.org directly. 8 do not, and of those 8, seven contain no link to any METR-controlled domain at all. One reaches metr.github.io. So seven chains I followed terminate somewhere that is not METR.

Four of the 25 print a figure of 14.5 hours (or "14 hours and 30 minutes") for Claude Opus 4.6. METR's own file says 718.8 minutes, which is 11h 59m. Three of those four assert 14.5 hours as current fact. The fourth, an AI 2027 tracker page, flags it as superseded and prints the corrected pair itself.

The clearest example of the qualifier evaporating is a beehiiv newsletter that renders it as: "Then Opus 4.6 dropped. 14.5 hours. That's not a typo. Nearly two full work days of expert-level labor." A coin flip has become replaced labor in the space of three sentences. AgentMarketCap keeps the "50%-time horizon" wording and still uses the wrong number.

Which doubling time is the real one?

There isn't one, and this is the part that gets flattened hardest. Time Horizon 1.1's own comparison table publishes three windows at once: 196.5 days all-time (stitched, using TH1 estimates for the pre-2023 models), 130.8 days [107, 161] since 2023, and 88.6 days since 2024. The May 2026 dashboard file publishes two more current fits: 187.8 days all-time and 128.7 days [104.4, 158.0] since 2023.

Seven months, four months and three months are all "METR's doubling time." Which one you quote is a choice about the start date, and the choice is the argument. Quoting one as "the" doubling time and then extrapolating it forward is where the forecasting posts go wrong.

The circulating "105 days" is not any of the five. The "240x jump" divides a 16-hour figure by a 4-minute figure; METR's file has Mythos Preview at 17h 25m and GPT-4o at 7.0 minutes, which is about 149x. Still enormous. Still not 240x, and not built from METR's numbers.

What does METR itself warn about?

METR hedges its own work harder than its citers do, in language worth quoting directly. On the confidence intervals: "These confidence intervals are still very wide." On the trend: "The trend in time horizon is somewhat sensitive to task composition." And on the foundation of the long-task measurements: "we measured human baseline times for only 5 of our 31 long (8h+) tasks. The remainder use estimated times."

Read that last one again. The 8-hour-plus end of the curve, the part every "agents can work a full day" headline rests on, has five measured human baselines behind it. The rest are estimates.

The dashboard adds a hard ceiling in a caption: "Measurements above 16 hrs are unreliable with our current task suite." METR's own top-scoring model, Mythos Preview at 17h 25m, is above that line. The YAML goes further and excludes it from the fit outright: the comment on the doubling-time block reads "excludes points with central estimate p50 > 16 hrs." METR does not use its own headline number in its own trend calculation.

What would change my mind?

If METR lands more human-baselined long tasks and the 80% horizon starts climbing on the same slope as the 50% horizon, the "agents can do a workday" reading becomes fair, because the coin flip stops being the load-bearing assumption. Right now the two curves are not the same shape, and the reliability people actually need for unattended work is the one growing slower.

The underlying finding also survives all of this. Task length that models can handle is growing fast on any of the five doubling fits, and that is a real, replicated, independently interesting result. The problem is not the research. It is that a measurement with an explicit threshold, an explicit task suite and an explicit "unreliable above here" line arrives at the reader as a workday. The same laundering happens to benchmark scores rebranded as "PhD-level reasoning", and it is why agent demos rarely survive contact with real work. The sibling problem, a benchmark score becoming a claim about whole jobs, is what GDPval numbers do to "matches human experts".

Verdict: 70% noise / 30% signal. The curve is real and steep. The sentence built on it, "agents can now work 14 hours autonomously," gets the number wrong, drops the success rate, ignores the task suite, and picks one of five doubling times without saying so. Read the threshold. Then read the second threshold.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗