Research
Research

Seven Vendors Published an OSWorld Score This Year. Not One Named the Human Baseline It Beat.

Aug 15, 2026 · 7 min read · by Jordan Kwan

TL;DR: Not as published, no. On 2026-08-15 I read 11 frontier vendor launch pages from 2026 for OSWorld claims. Seven name a score. Six name the version. Zero state which human baseline they are being measured against. The 72.36% figure everyone cites comes from the human study in the original 2024 OSWorld paper, run on a 369-task set, and vendors now report against OSWorld 2.0, a different 108-task benchmark whose own maintainers publish a best binary completion rate of 20.6%.

Start with the number itself, because almost nobody quoting it has read where it comes from. The OSWorld paper says: "While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success." That was 2024, on 369 desktop tasks, and it is a real human study, not a guess.

Two years later, vendors publish scores above it constantly, and the sentence "beats the human baseline" writes itself. The question is whether the two numbers are measuring the same thing.

What is on the official leaderboard, and how do you read it?

The canonical site is osworld-v1.xlang.ai. Note that os-world.github.io, the URL Anthropic still links to, 301s there; I confirmed the redirect with curl on 2026-08-15.

The leaderboard renders client-side and returns "Loading verified benchmark data..." to a raw fetch, so the rendered table is not readable by a fetcher. It is backed by two spreadsheets served from the same host, static/data/osworld_verified_results.xlsx and static/data/self_reported_results.xlsx, which I downloaded and parsed. Everything below comes from those files, not from a rendered page I could not see.

The site is explicit about what the verified sheet contains: "These are official results evaluated by our team under unified settings and environment." Self-reported results live in a separate file behind a separate link. That separation is the whole ballgame, and it is the benchmark maintainers doing the separating, not me.

The verified workbook holds 66 distinct entries. Fifteen sit above 72.36%. Fourteen of those fifteen are dated 2026:

Entry Score Type Official run dated
Intelligence-Indeed Agent 90.19 Agentic framework 2026-07-25
claude-fable-5 85.96 General model 2026-08-01
Pointer Agent w/ Opus 4.7 83.64 Agentic framework 2026-05-21
claude-opus-5 83.39 General model 2026-08-01
Muse Spark 1.1 80.67 General model 2026-07-09
MiniMax M3 75.19 General model 2026-06-08
Qwen 3.7 Plus 73.30 General model 2026-05-25
Kimi K2.6 73.06 General model 2026-04-20

So on OSWorld-Verified, which the maintainers describe as an in-place upgrade of the original 369-task set, models genuinely do clear 72.36%, and the benchmark team ran those numbers itself. That part of the story holds up better than I expected.

What did the vendor pages actually claim?

I fetched 11 launch posts and official model pages published in 2026 and grepped each for OSWorld. Seven name one:

Page Claim On official leaderboard Version named Baseline stated Labelled self-reported
Anthropic Sonnet 4.6 (Feb) OSWorld-Verified, chart image only Yes, 72.11 Yes No No
OpenAI GPT-5.5 (Apr) OSWorld-Verified 78.7% No Yes No No
Anthropic Opus 4.8 (May) OSWorld-Verified, Opus 4.7 restated to 82.3% No Yes No Partially
OpenAI GPT-5.6 (Jul) OSWorld 2.0 62.6% No Yes No No
Meta Muse Spark 1.1 (Jul) OSWorld named, no score on the blog Yes, 80.67 No No No
Google Gemini 3.6 Flash (Jul) OSWorld-2.0 33.8% No Yes No No
Anthropic Opus 5 (Jul) OSWorld 2.0, chart image only No, for 2.0 Yes No No

Four pages name no OSWorld figure at all: Anthropic's Opus 4.7 and Fable 5, and xAI's Grok 4.5 and 4.6.

Two of seven have the claimed model on the official verified board. Six of seven name the version, which is better discipline than I expected and worth crediting. Zero of seven state a human baseline. Zero label the figure self-reported.

The best-behaved page is Anthropic's Sonnet 4.6 launch, which footnotes its own chart: "Scores prior to Claude Sonnet 4.5 were measured on the original OSWorld; scores from Sonnet 4.5 onward use OSWorld-Verified," and says in the body that "the model certainly still lags behind the most skilled humans at using computers." That is a vendor pre-empting the exact error this post is about.

What is the spread on the same model?

Here is where it stops being a version-labelling quibble.

The OSWorld 2.0 team publishes its official results as JSON at osworld-v2.xlang.ai/static/data/leaderboard/official-results.json, last updated 2026-06-25, task version v2026.06.24, 108 tasks. Its top row is Claude Opus 4.8 at max reasoning with batched tool calls, 500 steps: 20.6% binary accuracy, 54.8% partial score. GPT-5.5 sits at 13.0% binary, 49.5% partial. The OSWorld 2.0 abstract states the same thing in words.

Now open OpenAI's GPT-5.6 table. Its OSWorld 2.0 row lists Claude Opus 4.8 at 54.8%.

Same model, same benchmark version, and the vendor number matches the official partial score to the decimal while the official binary completion rate for that model is 20.6%. The spread between the vendor figure and the benchmark team's completion figure for the same model is 34.2 points, and it is entirely a metric substitution. GPT-5.6 Sol's headline 62.6% is a partial-credit score. OpenAI never says so, and the row is captioned as a state of the art.

The cross-check confirms it. OpenAI lists GPT-5.5 at 47.5%; the official partial score for GPT-5.5 is 46.7% at 150 steps and 49.5% at 300. Its binary rate is 13.0%. Google's 33.8% for Gemini 3.6 Flash carries no metric label either, and I cannot tell you which of the two it is.

Does any of this mean the models got worse?

No. Computer use improved enormously and the verified board proves it: Anthropic's own Sonnet went from 43.90 in July 2025 to 72.11 in March 2026 under the same team's evaluation. Contamination is not the issue here, and no vendor in my sample published a number I can show to be wrong.

What I can show is that the sentence built on top of the number does not survive contact with the source. "Beats the human baseline" is doing three quiet substitutions at once: a partial-credit score standing in for task completion, a 2024 human study standing in for a task set that did not exist then, and the vendor's run standing in for the maintainer's. On OSWorld 2.0, the honest sentence is that the best officially evaluated agent finishes about one task in five, on workflows that take a skilled human a median of 1.6 hours.

I found no published human baseline for OSWorld 2.0. If one exists I have not seen it, and neither, apparently, has anyone comparing 62.6% to 72.36%.

What would change my mind?

One vendor printing the metric name next to the number. "62.6% partial score, OSWorld 2.0, self-run, 500 steps" would cost eight words and end the ambiguity permanently. Until then this belongs in the same bucket as every other launch-table figure nobody outside the lab has reproduced, and it is the same mechanism as the nine-month median age of the benchmarks in those tables: the measurement is fine, the sentence wrapped around it is not. The gap between a computer-use score and a computer-use agent is the same gap as every agent demo that works until you give it real work.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗