中文EN

The Token Economy Collapse: From Pay-Per-Use to Unlimited Subscriptions — Who Pays the Subsidy

Industry Signals · 2026-05-03

"Unlimited subscription" is the strongest demand signal of 2026 and the most anti-economic posture a supplier can take. Historically, every SaaS product with the word "unlimited" on it is a bloody internal bet. Somebody always eats the loss.

18 Months of Price Changes

Date Mainstream pricing
2024-Q1 Per-token, $3-$15/M tokens
2024-Q4 $20/month including X prompts
2025-Q2 Claude Pro $20, Claude Pro Max $200, with 5h/7d rate windows
2025-Q4 Codex Pro $200 unlimited, Cursor Pro $40 with heavily raised rate limits
2026-Q1 "unlimited" everywhere + 5h/7d soft quotas as the backstop

The "5h window" is the key design in this game — unlimited on the surface, in practice unlimited but throttled. It swaps pay-per-volume for pay-per-intensity.

Who's Paying for "Unlimited"

Loser 1: Anthropic / OpenAI (the frontier labs)

Unlimited is a win for consumer psychology. On the P&L:

  • The top 1% of heavy users consume 100x what a normal user does
  • If a heavy user pays $200 and a normal user also pays $200, the P&L model is really "normal users subsidize heavy users"
  • And heavy users keep using more — marginal cost is 0, after all

Their options here: keep burning investor money, tighten rate limits (already happening), or cut prices while strictly capping volume (impossible — the narrative would collapse).

Loser 2: mid-tier tooling vendors (Cursor / Windsurf / Cline and friends)

The Cursor play — an IDE layer on top of somebody else's model — used to make money on a 30-50% token markup. In an unlimited world, what do you charge for?

Their responses:

  • train their own model (money into a black hole)
  • push enterprise ACV (competing on sales, not product)
  • pivot to an MCP platform (covered in the previous post)

But the core problem is unsolved: why wouldn't a user just use the official Claude Code / Codex CLI, which is also unlimited?

Loser 3: heavy users themselves

Counterintuitive, but true.

  • Rate limit quota is the new friction — with 10% left in your 5h window you don't dare kick off a big task. That's a mental tax.
  • The silent downgrade trap — unlimited looks infinite, but hitting the ceiling force-downgrades you to a mini model, and users often don't notice.
  • Your data accretes to the vendor — the more you run, the better they understand your working patterns, the more precisely they can price you later.

Heavy users got unlimited and lost token-by-token controllability.

Loser 4: normal users (the most invisible one)

A normal user consumes less than $5 of tokens a month and is required to pay $20. That's the most invisible subsidy in the system. "Unlimited" pulls in 5x as many light users to foot the bill.

How This Game Ends

SaaS history has run the "unlimited" play and corrected it over and over. The correction path is always the same:

Phase A: open unlimited to grab users (where we are now)

Phase B: quietly add quotas and fine print (already started: 5h, weekly caps, monthly soft limits)

Phase C: tiering, splitting the high-end users out (coming: $500/month, $1000/month enterprise plans)

Phase D: "fair use" gets spelled out and unlimited becomes a marketing word (my guess: mid-2027)

The vendor that admits this first wins. Getting users to accept "unlimited is an abstraction, the real quota is XX" in the narrative beats quietly choking them.

How to Position Your Own Budget

If you're an individual developer:

  1. This is a historical window — unlimited really is (mostly) unlimited right now
  2. It gets quietly tightened starting next year — don't build the habit of depending heavily on one vendor
  3. Keep a local model as the backstop — the posture from the previous post
  4. Check your actual consumption once a month — if you've tripped the rate limit more than 3 times, you're the one being subsidized. Enjoy it, no shame in that. Just know the arbitrage window closes.

If you run a SaaS company:

  1. Don't do unlimited yourself — this is the labs' capital game and you can't fund it
  2. Price on outcomes — $X per successful review, $X per bug fixed
  3. Pass token cost straight through — tell the customer "this task used $0.5 of GPT-5.5 + $0.05 of Qwen-Max," let them see the value
  4. Stock up on cheap model alternatives — Qwen / DeepSeek for tasks that don't need frontier quality; spend the savings back on the customer

A Counterintuitive Prediction

The "unlimited" era ends within 12 months. We go back to outcome-based pricing — not the crude token-by-token metering, but pricing by completed work.

  • One PR review = $0.50
  • One bug localized = $1.00
  • One feature spec → implementation = $20

That's the pricing that actually makes sense in the AI era — tied to value, not to token counts.

The token economy collapse is inevitable, and the next pricing paradigm is being born right now. Whoever moves first eats this wave.

← More in Industry Signals