Research
Research

I Dated Every Benchmark in Five 2026 Launch Tables. The Median Is Nine Months Old.

Aug 15, 2026 · 7 min read · by Jordan Kwan

TL;DR: Because the old benchmarks got contaminated, and the replacements are too new for anyone outside the lab to have checked them. On 2026-08-15 I enumerated every benchmark named in the headline score tables of five frontier launches published since 2026-06-01. That is 54 distinct benchmarks. I could find a dated public artifact for 25, and their median age on launch day was 9.1 months. The remaining 29 had no dated public artifact I could find. Only 11 of the 54 can be evidenced as having existed 12 months before the launch that cited them.

On 23 February 2026, OpenAI published a post explaining why it stopped reporting SWE-bench Verified. It is worth reading in full, because the honest version of the finding is stronger than the version circulating.

What did OpenAI actually say?

Two things, and people keep merging them.

The first is about test quality. OpenAI audited 138 problems, 27.6% of the 500-problem set, specifically chosen because o3 "did not consistently solve" them over 64 runs. In that adversarially selected subset, "at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions." That 59.4% is a defect rate inside a deliberately hard slice. It is not a contamination percentage, and anyone quoting it as one has not opened the post.

The second is about training exposure, and it is the part that ends the benchmark. OpenAI built a red-teaming setup where GPT-5 probed GPT-5.2-Chat, Claude Opus 4.5 and Gemini 3 Flash Preview over 15 turns per task. It found "all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks." Opus recalled an inline code comment word for word. Gemini 3 Flash reproduced task text given only an ID.

OpenAI's conclusion: "improvements on SWE-bench Verified no longer reflect meaningful improvements in models' real-world software development abilities." So they stopped reporting it, and recommended everyone else do the same.

That is the correct call. A leaked benchmark is a broken instrument, and retiring it is what a serious lab does. Which is exactly why the churn deserves measuring rather than sneering at.

What did I count?

I took every frontier launch post published between 2026-06-01 and 2026-08-15 and enumerated every benchmark named in its headline score table. Five had tables I could read cell by cell:

Launch Date Benchmarks in the table
OpenAI GPT-5.6 2026-07-09 34
xAI Grok 4.5 2026-07-16 5
Google Gemini 3.6 Flash 2026-07-21 18
Anthropic Claude Opus 5 2026-07-24 11
xAI Grok 4.6 2026-08-12 10

Two more launches in the window resisted enumeration. Anthropic's Claude Fable 5 announcement says "the table below compares the capabilities of Fable 5 and Mythos 5 to other leading models," and that table ships as a bitmap. Meta's Muse Spark 1.1 blog carries no scores and points to a PDF evaluation report my parser could not read. I am reporting five of seven, not seven of seven.

Deduplicated, the five tables name 54 distinct benchmarks. For each I recorded three things: whether a dated public artifact exists from at least 12 months before the launch naming it, whether a public leaderboard exists that I could open, and whether the page states the run was executed by someone other than the vendor.

For dates I used the arXiv API (first submission date of the v1 abstract) and the GitHub API (created_at on the benchmark's repository). Both are reproducible in one request each.

How old is a benchmark in a 2026 launch table?

I could date 25 of the 54. Their median age at the launch that cited them was 9.1 months. The distribution is what makes the number interesting rather than the number itself.

The eight youngest, with the evidence:

Benchmark First dated artifact Age at launch
DeepSWE v1.1 arXiv 2607.07946, 2026-07-08 1 day
OSWorld 2.0 osworld-v1.xlang.ai news line, 2026-06-26 0.4 months
Agents' Last Exam arXiv 2606.05405, 2026-06-03 1.2 months
SWE-Marathon arXiv 2606.07682, 2026-06-05 1.3 months
SEC-bench Pro arXiv 2605.26548, 2026-05-26 1.4 months
ExploitGym arXiv 2605.11086, 2026-05-11 1.9 months
BenchCAD arXiv 2605.10865, 2026-05-11 1.9 months
AutomationBench GitHub Zapier/automationbench created 2026-02-23 4.5 months

DeepSWE is the one to sit with. Its paper went up on arXiv on 8 July 2026. OpenAI's GPT-5.6 launch, which reports a DeepSWE v1.1 score, went up on 9 July. Grok 4.5, Grok 4.6 and Gemini 3.6 Flash all report it too. Four labs converged on a benchmark that had been public for a day.

The other 29 benchmarks are the real story. I searched arXiv and GitHub for each and found no dated public artifact: CursorBench, FrontierCode 1.1, Frontier-Bench v0.1, AA-Briefcase, GeneBench Pro, LifeSciBench, BioMysteryBench, Toolathlon, GraphWalks, gdp.pdf and 19 others. Some are explicitly vendor-internal and labelled as such, which is at least honest: OpenAI marks "Management Consulting Tasks (Internal)" and "MedChemBench (Internal)" in its own table. If you count those 29 as zero months old, because launch day is the earliest a reader could have seen them, the median age of the set drops to zero.

Only 11 of 54 have a dated artifact old enough to have existed a year before the launch that cited them: GPQA Diamond, CharXiv, LVBench, MMMU-Pro, FrontierMath, Humanity's Last Exam, BrowseComp, HealthBench, OSS-Fuzz, and both Terminal-Bench rows, 2.1 and 3.0, which trace to one repository created 2025-01-17 despite the version numbers. Absent: SWE-Bench Pro, repo created 2025-09-05, and OSWorld 2.0, younger than every launch citing it.

Can you check any of these numbers yourself?

Rarely. I looked for a public leaderboard for each of the 54 and found one for 16. The ones that resolved on 2026-08-15 include tbench.ai for Terminal-Bench, Scale's SWE-Bench Pro board, ARC Prize, Epoch's FrontierMath page, Vals AI and the OSWorld 2.0 leaderboard. Thirty-eight have no board at all, which means the vendor's number is the only number, and the boards that do exist do not converge: on one day, seven of them named six different models at rank 1.

Eleven of 54 state a third-party run. xAI is the surprise here. Its Grok 4.5 chart caption reads "Eval created by Datacurve, run with each model provider's harnesses by AA," and its DeepSWE 1.1 figures are labelled "mini-swe-agent harness run by Datacurve." Anthropic footnotes Opus 5's Frontier-Bench figure as "an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task." Both name the harness. Neither is unreproducible in principle. What is unavailable is anyone else's run of the same thing, which is how only 2 of 17 Humanity's Last Exam claims match the board they are attributed to.

So is the churn a scam?

No, and treating it as one gets the diagnosis backwards. Retiring a memorized benchmark is the responsible move, and the labs doing it are publishing the evidence that forces their own hand. OpenAI recommended competitors drop a benchmark OpenAI itself created.

The cost is separate from the motive. A benchmark accumulates external checkability slowly: independent runs, a leaderboard, a replication that disagrees. Replacing your evals every nine months means the score table is permanently made of instruments too young to have any of that. Contamination made the old numbers meaningless. Novelty makes the new numbers unfalsifiable. Both can be true, and right now both are.

This is the same disclosure gap I found when I read 23 model launches looking for a checkable number, measured from the other end: not what the page omits, but how old the thing being measured is. It also drives the specific failure in what "beats the human baseline" means on a computer-use benchmark, where a 2024 human number gets attached to a 2026 task set.

What would change my mind?

One benchmark surviving three consecutive flagship generations with a public leaderboard and at least one independent run per model. That is a low bar and nothing in my sample clears it. Until something does, treat a launch table the way you would treat any expert-level claim with no working shown: probably directionally right, and nine months from having any outside evidence behind it.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗