Research
Research

Only 3 of 21 Model Launches Still Report SWE-bench Verified. 12 Switched to the Benchmark OpenAI Retracted in July.

Aug 15, 2026 · 7 min read · by Jordan Kwan

TL;DR: Mostly no. On 2026-08-15 I read every frontier model launch page and official model card I could reach that was published between OpenAI's 2026-02-23 warning and today. 21 documents across 9 labs resolved. Of those 21, only 3 still report SWE-bench Verified, and all 3 report SWE-bench Pro in the same table. None lead with Verified. 12 of 21 report SWE-bench Pro, which is the awkward part: OpenAI retracted its own recommendation of SWE-bench Pro on 2026-07-08 after estimating that roughly 30% of its tasks are broken.

The premise I started with was that vendors ignore inconvenient warnings. The count says the opposite happened, and something stranger took its place.

What did OpenAI actually say?

On 23 February 2026, OpenAI published Why SWE-bench Verified no longer measures frontier coding capabilities. Two findings, and they are not the same finding.

First, test quality. OpenAI audited 138 Verified problems that OpenAI o3 failed to solve consistently over 64 runs, had each reviewed by at least six engineers, and found "59.4% of the 138 problems contained material issues in test design and/or problem description." That number is about a hand-picked hard subset, not about the 500-problem set, and it is not a contamination rate. The circulating version of this stat treats 59.4% as "SWE-bench Verified is 59% contaminated." It is not. It is the share of already-failing problems whose tests were broken.

Second, contamination, measured separately. OpenAI ran a red-teaming setup in which GPT-5 probed GPT-5.2-Chat, Claude Opus 4.5 and Gemini 3 Flash Preview over 15 turns, and reported that "all frontier models we tested were able to reproduce the original, human-written bug fix used as the ground-truth reference, known as the gold patch, or verbatim problem statement specifics for certain tasks." The published transcripts show Opus quoting an inline code comment word for word. OpenAI's conclusion: "This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too." Its replacement suggestion was explicit: "OpenAI recommends reporting results for SWE-bench Pro."

What exactly did I count?

I fetched every frontier launch announcement or official model card published in the window 2026-02-23 to 2026-08-15 that I could read as a primary document, and scored each on three states: reports Verified, reports Pro, reports neither. Twelve labs were attempted. Alibaba's Qwen cards returned 401 and qwen.ai/blog rendered no text; Mistral's release page 404'd; Microsoft's MAI post 404'd. Those three are absent from the denominator rather than guessed at.

One scoping note that matters. I read announcement pages and model cards, which is where the marketing number lives. I did not read every system card, and OpenAI's own statement covers its blog and eval reporting. A lab can drop a benchmark from its launch page and keep it in a 90-page appendix.

Launch (doc read) Date Verified Pro
OpenAI GPT-5.4 2026-03-03 N Y
Google Gemini 3.1 Flash-Lite (card) 2026-03-03 N N
xAI Grok 4.2 (docs) 2026-03-09 N N
OpenAI GPT-5.4 mini/nano 2026-03-17 N Y
MiniMax M2.7 (card) 2026-03-18 N Y
Z.ai GLM-5.1 (card) 2026-04-06 N Y
Anthropic Opus 4.7 2026-04-16 Y Y
xAI Grok 4.3 (docs) 2026-04-17 N N
Moonshot Kimi K2.6 2026-04-20 Y Y
OpenAI GPT-5.5 2026-04-23 N Y
DeepSeek V4-Pro (card) 2026-04-24 Y Y
Google Gemini 3.5 Flash (card) 2026-05-20 N Y
Anthropic Opus 4.8 2026-05-30 N N
Anthropic Fable 5 / Mythos 5 2026-06-09 N N
Z.ai GLM-5.2 (card) 2026-06-16 N Y
OpenAI GPT-5.6 2026-07-09 N Y
Meta Muse Spark 1.1 2026-07-09 N N
Moonshot Kimi K3 (card) 2026-07-16 N N
xAI Grok 4.5 2026-07-16 N Y
Anthropic Opus 5 2026-07-24 N N
xAI Grok 4.6 2026-08-12 N N

Who still reports SWE-bench Verified?

Three documents: Anthropic's Opus 4.7, Moonshot's Kimi K2.6 and DeepSeek's V4-Pro card. All three publish Pro in the same table, so none of them is defending Verified as the headline. Anthropic went further and footnoted it: the Opus 4.7 announcement states "Our memorization screens flag a subset of problems in these SWE-bench evals. Excluding any problems that show signs of memorization, Opus 4.7's margin of improvement over Opus 4.6 holds." That is a lab conceding the contamination and then showing the result survives it.

Anthropic then stopped entirely. Opus 4.8, Fable 5 and Opus 5 name no SWE-bench variant at all, reporting Frontier-Bench, CursorBench, Terminal-Bench 2.1 and OSWorld instead. Between April and May, the most-quoted coding benchmark in AI vanished from Anthropic's launch pages without an announcement.

Where did everyone go instead?

To SWE-bench Pro, exactly as OpenAI suggested. 12 of 21 documents report it, including all four OpenAI launches, both Google model cards (the Gemini 3.5 Flash card gives it a "SWE-Bench Pro (Public)" row), both Z.ai cards, MiniMax, DeepSeek and xAI's Grok 4.5.

Then, on 8 July, OpenAI published Separating signal from noise in coding evaluations. Its automated pipeline flagged 200 broken tasks (27.4%) in the 731-task public split; its human annotation campaign, five engineers per task, found 249 (34.1%). The line that ends the story: "Given the issues uncovered in this analysis, we retract our earlier recommendation to adopt SWE-Bench Pro." The same post notes frontier models went from 23.3% to 80.3% on that split in eight months, which is a suspicious slope for a benchmark built to resist saturation.

Did anything change after the retraction?

Barely, and the exception is funny. Five launch documents in my sample were published after 8 July. Two still report SWE-bench Pro. One of them is OpenAI's own GPT-5.6 page, published 9 July, one day after OpenAI retracted the recommendation, carrying a SWE-Bench Pro row in its comparison table. The other is xAI's Grok 4.5, whose page also states that "Competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards," meaning it is comparing its own Pro runs against numbers it did not reproduce on a benchmark its originator now says is a third broken.

The three launches after 8 July that report neither (Kimi K3, Opus 5, Grok 4.6) did not do so in response to anything. They had already moved to Terminal-Bench, DeepSWE, FrontierCode and internal evals months earlier.

What does this not prove?

It does not prove SWE-bench Verified is dead in the wild. My denominator is lab launch documents, not coding-tool marketing pages, not comparison listicles, not system cards. Verified could easily still be the default line in vendor sales decks, and 21 documents cannot tell you that.

It also does not prove the labs deferred to OpenAI. Most of these releases were probably already migrating off a saturated benchmark for their own reasons, and OpenAI's post landed on a door that was already open.

What it does show is a measurement layer with no floor. The industry standard was retired in February. Its officially recommended replacement was retracted in July, by the same lab, using the same audit method. Nine of 21 launches now report no SWE-bench variant at all and substitute internal evals nobody outside the company can run. That is the pattern worth watching, because it is the same one behind every launch table that publishes scores and no outputs and behind HLE numbers that appear on no leaderboard that ran them. Benchmarks keep getting retired faster than replacements get trusted, and the fallback is always the vendor's own unpublished ruler: across five 2026 launch tables, the median benchmark was 9.1 months old and 29 of 54 had no dated public artifact at all. We have been here before with coding productivity studies, and it ended the same way.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗