Token Prices Fell 280x. Your AI Bill Went Up. Both Are True.
Aug 6, 2026 · 5 min read · by Jordan Kwan
TL;DR: The cost of querying a GPT-3.5-class model fell more than 280-fold in eighteen months, and yet I could not find a single AI product that loosened its pricing in the last fourteen: my compiled list of documented changes runs seven for seven in the tightening direction. The resolution to the paradox is that you do not buy tokens, you buy tasks, and reasoning models plus agents made tasks roughly an order of magnitude hungrier while vendors stopped subsidizing the difference.
"AI is getting expensive" and "AI is getting cheaper" are both load-bearing claims in 2026, usually deployed by people selling something. So I built the record: every pricing or limit change I could verify against a primary vendor source, June 2025 through August 2026.
What did I count?
Seven documented changes, seven products, one direction:
| Product | When | What changed |
|---|---|---|
| Cursor Pro | Jun 2025 | 500 fast requests became a $20 usage credit at API rates; CEO apology and refunds followed |
| Replit Agent | Jun 2025 | flat $0.25 checkpoints became variable "effort-based" pricing |
| Claude Pro/Max | Aug 2025 | weekly rate limits stacked on the 5-hour window |
| Salesforce | Aug 2025 | ~6% average list increase, framed around AI innovation |
| Canva Pro | 2025 | $12.99 to $15, linked to AI features (price trackers; no primary announcement) |
| GitHub Copilot | Jun 2026 | premium requests became token-metered AI Credits at API rates |
| Claude Sonnet 5 API | Sep 2026 | introductory $2/$10 per million tokens rises to the standard $3/$15 |
Six of seven rest on primary vendor announcements; Canva rests on price trackers. Zero documented cases of an AI product loosening limits or cutting a subscription price in the same window. The pattern within the pattern: three of the seven (Cursor, Replit, Copilot) are the same event, a flat rate converting to metered billing after the vendor discovered what its heaviest users actually cost.
Why are bills rising while token prices collapse?
Because the unit that got cheap is not the unit you consume. Stanford's AI Index documents the collapse: querying a model at GPT-3.5's level fell from $20 per million tokens in late 2022 to $0.07 by late 2024, and Epoch AI finds inference prices falling 9x to 900x per year depending on the task. But over the same period, the products moved to reasoning models that consume roughly an order of magnitude more tokens per task, and then to agents that run for hours. A cheap token multiplied by ten times the tokens, running unsupervised overnight, is how Cursor's "unlimited" plan died and why every premium tier now meters something fuzzy in units that do not survive comparison: 25 pricing pages carry 18 distinct names for the billing unit, and a "credit" worth a dollar at one vendor and a cent at another.
GitHub's chief product officer Mario Rodriguez said the honest version out loud when Copilot switched: usage-based billing "better aligns pricing with actual usage, helps us maintain long-term service reliability, and reduces the need to gate heavy users." Translation: the flat rate was a promotional artifact, and the promotion is ending. Developer reaction, as Visual Studio Magazine's headline compressed it: "You will get less, but pay the same price."
What does heavy real-world usage actually cost?
Ed Zitron's reporting supplies the enterprise receipts: Zillow spent $749,000 on tokens in April 2026 alone, on pace for seven figures a quarter, and Stripe's technical staff were running roughly $94,000 a day. His core line is the thesis of the whole cost story: "LLMs are too expensive! They cost too much to run, and said costs appear to increase linearly with revenues." And the vendors are on the wrong side of the same math at incomprehensible scale: OpenAI targeted over $20 billion in annualized revenue against about $1.4 trillion in compute commitments, by Sam Altman's own numbers, with audited 2025 losses reported at roughly $38.5 billion. That revenue figure is also the last one OpenAI said on the record: of 24 traced AI revenue numbers, eight came from the companies and none from a filing. A vendor losing billions on your subscription eventually reprices your subscription. Every row in my table is that sentence happening.
What does this not prove?
It does not prove per-task costs will only rise: the efficiency curve is real too, and the ARC-AGI benchmark's cost per task fell from roughly $4,500 with o3 in late 2024 to about $11.64 with GPT-5.2 a year later, a ~390x gain. When capability targets hold still, costs crater. The ratchet tightens because targets never hold still; each new model generation spends its efficiency dividend on more reasoning, more steps, more autonomy. It also does not prove vendors are gouging: Cursor's apology and Anthropic's under-5% estimate both read like companies genuinely surprised by their own tails. And it does not settle the bigger solvency question, whether an industry losing this much money on every subscriber can keep existing at current prices, which is the bubble debate's home turf rather than this post's.
What should you do about it?
Three operational conclusions from the table. Budget per task, not per seat: any flat-rate AI line item in your budget is a number the vendor can convert to metered billing with thirty days' notice, because seven of seven did. Watch introductory pricing expiry dates the way you watch contract renewals; Sonnet 5's September step-up is on a published calendar, and it will not be the last. And instrument your own usage now, before the next conversion: the difference between the Zillow bill and a sane one is knowing which tasks are worth ten thousand reasoning tokens, and nobody's pricing page will figure that out for you.
Written by Jordan Kwan, founder of Reachium.
I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.
See what Reachium does ↗