Tool Teardowns
Teardown

I Traced 17 Humanity's Last Exam Scores. Two of Them Are on a Leaderboard That Ran the Test.

Aug 15, 2026 · 6 min read · by Jordan Kwan

TL;DR: Usually from a different run than the one you think. On 2026-08-15 I took 17 Humanity's Last Exam score claims, 10 from leaderboard sites and 7 from vendor launch documents, and checked each against the live CAIS board, the static table on agi.safe.ai, Scale's SEAL leaderboard and Artificial Analysis. Only 2 of the 17 match a number on an official CAIS or Scale surface. Most of the gap is legitimate: tools-enabled runs score 12 to 19 points higher than text-only runs of the same model. The problem is that all 7 vendor claims state which run they are quoting and 0 of the 10 aggregator claims do.

Where is the official leaderboard, and is there more than one?

Two, on the same domain, and the difference matters before you accuse anyone of anything.

Open agi.safe.ai and you get a static results table under the heading "Quantitative Results," captioned "Judge Model: o3-mini | Dataset Updated: April 3rd, 2025." Ten models. The leader is Gemini 3 Pro at 38.3%, with GPT-5 at 25.3% and Grok 4 at 24.5%. Nothing released in 2026 appears.

The same page also renders a live board, which arrives client-side and shows nothing to a raw fetch. Its data endpoint is dashboard.safe.ai/api/models, which served me 53 models with HLE scores when I sent it an Origin: https://agi.safe.ai header. That board is current: it carries Claude Opus 5 (released 24 July) at 51.0, Claude Fable 5 at 52.72, Muse Spark 1.1 at 49.44 and GPT-5.6 Sol at 45.52.

So one page shows two leaders. If you fetch it the naive way, you conclude HLE was abandoned sixteen months ago. It was not. The static table is a frozen artifact of the paper, and its numbers agree with the live board for every model it contains.

The third surface is Scale's SEAL leaderboard, which runs its own evaluation at temperature 0 with o3-mini as judge and publishes confidence intervals. It tops out at gemini-3.1-pro-preview (thinking high) at 46.44 ± 1.96, followed by gpt-5.4-pro at 44.32 and Muse Spark at 40.56. It lists no model released after roughly May 2026. Different cadence, not abandonment.

What exactly did I check?

I searched "Humanity's Last Exam leaderboard," took the leaderboard pages from the results, and recorded the model-and-score claims each one puts at the top. That gave 10 claims. I added 7 HLE numbers published by vendors in launch pages and model cards. Then I checked all 17 against the live CAIS board, the static table, SEAL, and Artificial Analysis, whose page carries 567 model-configuration rows because it runs each model at several reasoning-effort levels.

One page could not be read. llm-stats.com returns no scores in its raw HTML because its board renders client-side, so I recorded it as unread rather than as empty.

How many of the quoted scores are on an official board?

Two. Both are from llmrun.dev, which says plainly "llmrun does not run this benchmark," credits Epoch, and stamps its data "through Apr 2026." Its Gemini 3.1 Pro Preview figure of 46.4% is SEAL's 46.44. Its Kimi K2.5 figure of 24.4% is SEAL's 24.37. Cited, dated, checkable, correct.

Three more are traceable one step further out. pricepertoken.com reports Claude Fable 5 at 55.5%, Claude Opus 5 at 54.9% and GPT-5.6 Sol at 49.5%, and prints "Source: Artificial Analysis." All three match Artificial Analysis exactly. Not an official board, but an independent runner named on the page, which is the whole ask.

That leaves 12 of 17 that match nothing I could find on CAIS, SEAL or Artificial Analysis.

Does the tools split explain the gap?

For the vendor numbers, largely yes, and this is the part the "benchmarks are fake" version of this story gets wrong.

OpenAI's GPT-5.4 mini and nano page publishes two separate rows: "HLE w/ tool" at 52.1% and "HLE w/o tools" at 39.8%. Same model, same day, 12.3 points apart. Moonshot's Kimi K2.6 post is more explicit still: on the text-only subset it "achieves 36.4% accuracy without tools and 55.5% with tools," a 19.1-point gap, and it documents the tools used (search, code interpreter, web browsing) and the 262,144-token generation limit for the tools run. Google's Gemini 3.5 Flash card labels its 40.2% as "full set, text + MM" with "No tools," and the Gemini 3.1 Flash-Lite card labels its 16.0% the same way.

That is 7 vendor claims out of 7 stating the condition. Two numbers 15 points apart for one model can both be honest. The official boards run text-only without tools, which is why a vendor's tools-enabled figure legitimately will not appear there.

Zero of the 10 aggregator claims state the condition.

Which numbers do not survive any explanation?

benchlm.ai leads with "Claude Opus 5 leads the HLE leaderboard with 64.7%, followed by Claude Mythos 5 (64.5%) and Muse Spark 1.1 (62.1%)," dated 15 August 2026, across "50 tracked models."

CAIS's live board has Opus 5 at 51.0 and Muse Spark 1.1 at 49.44. Artificial Analysis's highest-effort Opus 5 run is 54.9 and its Muse Spark 1.1 run is 46.2. SEAL lists neither. Claude Mythos 5 appears on no board I checked. So 64.7% sits 13.7 points above the official board and 9.8 points above the best independent run, and tools cannot rescue it, because the same page tells you it should not: "A model with search, browsing, or code execution is not taking the same test as a closed-book model," and "Operator receipt: 50 sourced rows are currently displayable on this page." The visible rows carry one label, "Closed," which marks closed weights, not closed book.

aitooltier.com is the other one. It ranks "Muse Spark (Meta)" at 58% and "Grok 4.20" at 50.7%, citing "Official source https://lastexam.ai/." CAIS has Grok 4.2 at 30.2. The same page describes HLE as "a 3,000-question benchmark crowdsourced from thousands of subject-matter experts," while the source it links says 2,500 questions from nearly 1,000 contributors.

What would change my mind?

A protocol column. Not more leaderboards, and not a takedown of the ones that exist. If every quoted HLE number carried three fields, tools on or off, which subset, which harness, then 12 of my 17 orphans would probably resolve into ordinary run-to-run variation and the two genuinely invented ones would stand out immediately.

Until then, an HLE percentage with no condition attached is not a measurement, it is a mood. The vendors are doing this correctly and the layer that repackages them is not, which is the same failure I found when checking which model launches publish anything you can verify and when tracking where the labs went after SWE-bench Verified. Treat a bare HLE score exactly like a claim of PhD-level reasoning: ask which test, under what conditions, graded by whom.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗