Tool Teardowns
Teardown

ChatGPT Work Is a Month Old and Almost Nobody Has Actually Tested It

Aug 6, 2026 · 5 min read · by Jordan Kwan

TL;DR: ChatGPT Work, launched July 9, 2026, is OpenAI's replacement for the retired agent mode: an agent that "can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work." The early evidence pattern is consistent: bounded, file-shaped deliverables from connected data genuinely finish, in minutes; open-ended and local-file visual tasks still fail. But the loudest finding is the vacuum: a month in, the top of search contains exactly one named-author hands-on review.

When a product this big is this new, the honest move is to audit what is actually known rather than pretend to a month of testing nobody has had time for. So this is built from OpenAI's official launch materials and help docs, the predecessor's track record, the handful of real early reviews, and a count of the review economy itself.

What does OpenAI say it does?

Per the launch post, ChatGPT Work produces "finished materials like sheets, slides, docs, and web apps," runs scheduled and event-triggered tasks, and on desktop can operate the computer directly, "clicking, typing, and moving files." It ships with GPT-5.6 in three effort tiers named Sol, Terra, and Luna, connects to more than 1,400 plugins, and is available on every plan via the desktop app, with web and mobile for Plus and up. The oversight model is genuinely more built-out than agent mode's was: a Plan mode that proposes steps for approval, user-set checkpoints, and an "Auto-review" layer that OpenAI says uses its most advanced models to review important actions before they happen.

Note what is missing from all of it: a task quota you can plan around. There is no "X tasks per month" anywhere; usage draws from a metered agentic credit pool shared with Codex and "varies with the amount of work required." The retired agent mode had legible caps (40 messages a month on Plus, 400 on Pro). Its successor's unit of consumption is, like every premium AI tier this year, deliberately fuzzy: across 25 AI pricing pages there are 18 distinct names for the billing unit and four vendors who never state a rate anywhere public.

What happened to agent mode?

It is gone. OpenAI's help center now says plainly: "ChatGPT agent is no longer available. Use ChatGPT Work for longer, multi-step tasks and finished deliverables." That retirement matters for expectations, because agent mode's record is the best predictor we have. At launch it posted impressive benchmark numbers (41.6 on Humanity's Last Exam, state of the art on browsing), but its most instructive score was 45.5% on SpreadsheetBench against a 71.3 human baseline, and the most careful independent test, Timothy B. Lee's at Understanding AI, concluded: "the agent is nowhere close to the level of reliability required for me to really trust it," adding that "an agent that frequently does the wrong thing is often worse than useless." One year and one product generation later, that is the bar ChatGPT Work has to clear.

What does it actually finish?

The early pattern, from the few real tests that exist: bounded tasks with clear inputs and a file-shaped output genuinely complete, and fast. The synthesis of early testing at Smart AI Library reports clean output in one to nine minutes for report-building, site-generation, and data-pull tasks, while the failures cluster exactly where you would predict: a visual report from local images ran half an hour and produced nothing usable, and polished spreadsheets occasionally contained confidently wrong counts, the demo-to-production signature in miniature. The one named-author hands-on in the search top ten, Eric Bye's, credits it as able to "grind through huge amounts of context and complete genuinely complex tasks" while warning that "access alone will not create value" without training, cost controls, and governance.

Competitively, this is now a three-way race among incumbents' work agents. Anthropic's Claude Cowork, which "completes tasks you can steer from anywhere," went from desktop research preview in January to web and mobile by July, and Google's Gemini Spark launched at I/O in May as "a 24/7 personal AI agent" for its Ultra subscribers. All three sell the same promise, and all three share the same absence of a published error rate. The launch also quietly consolidated OpenAI's sprawl: the Codex app merged into the new desktop app, the old one was renamed ChatGPT Classic, and the standalone Atlas browser began sunsetting the same day, which tells you how central OpenAI considers this product.

Should you use it?

Yes, as a metered trial, with the meter watched, and the barrier to trying it is genuinely low: the desktop app includes Work on every plan, free included, so the experiment costs nothing but the credits. Give it the tasks the evidence supports: recurring, bounded, pulling from sources it can connect to, ending in a sheet, deck, doc, or simple app a human will inspect. Keep Plan mode on and checkpoints tight, and treat anything touching local files or visual judgment as not yet its job. And discount the content you will read about it: as my count shows, a month after launch the top ten results for its review query contain one actual test, while searches for its retired predecessor still surface reviews presenting agent mode as a current product. The review economy is slower than the release cycle now. For a tool sold on finishing your work, the most useful thing you can do is make it show its work first.

Written by Jordan Kwan, founder of Reachium.

I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.

See what Reachium does ↗