Imagine waiting for a big Tuesday announcement and getting a long usage-policy post first.
The headline was uncomfortable: the $200 Pro plan would reopen to new subscribers on September 30, 2026, but under the new accounting method its allowance would amount to roughly half the API-dollar spend of the old $200 plan.[1]
It feels like waiting for fireworks and receiving a revised utility bill first.
There is one important distinction, though. The announcement did not say that every model would simply get half as many messages. In the same week, OpenAI said GPT-6 Sol and Luna API prices were 50% lower than the promotional pricing of their GPT-5.6 predecessors.[2][3]
The vendor's argument is therefore straightforward:
Cut the API-dollar-equivalent allowance, lower model costs, and still increase the amount of work users can complete.
That arithmetic can work.
But customers do not buy "API-dollar equivalents." They buy completed work.
1. What exactly is being halved?
Tibo's September 29 post said the $200 Pro tier would reopen the next day and that, under the new usage calculation, it would net out to half the API spend of the previous $200 plan.[1]
The same post said the five-hour limit would not return, users could spend their weekly allowance when they wanted, future model-efficiency gains would be passed through in lower API prices, and additional subscription features that do not draw from usage would be announced.[1]
OpenAI's September 22 announcement stated that GPT-6 Sol and Luna have API prices 50% below GPT-5.6 promotional pricing.[2] The current pricing page confirms the new model prices.[3]
On paper, the math is clean.
Let the old allowance be B and the API-equivalent cost of one job be C. Old throughput is roughly B/C.
If the new allowance becomes 0.5B while the same job becomes 0.5C, then:
0.5B / 0.5C = B / C
Theoretical throughput stays constant.
Real AI workloads are not that uniform.
2. Users measure finished jobs, not abstract dollars
A hundred short questions and one repository-scale repair are not remotely equivalent.
Long context, deeper reasoning, tool calls, verification passes, retries, and agent loops can turn one request into a large unit of consumption.
That creates a very different user experience:
"I can still send messages" is not the same as "I can finish another serious task this week."
The useful metric for a heavy user is therefore not messages or tokens. It is closer to:
completed jobs / monthly subscription cost
If a model is brilliant but the allowance runs out halfway through the work, intelligence alone does not create a finished artifact.
3. "It feels like one-hundredth of Opus" is not a benchmark, but the underlying complaint is real
People doing heavy agentic work can experience enormous differences between subscription limits.
One service may allow several long jobs in succession. Another may feel as if one large task nearly empties the tank.
Saying "it feels like one-hundredth as much" is obviously not a measured ratio. It depends on task type, plan, model, context length, and reasoning intensity.
But dismissing the exaggeration misses the real issue: the granularity of the work must fit the granularity of the limit.
For five-minute tasks, a small allowance can still be useful.
For work that takes thirty minutes, an hour, or several repair-and-test loops, running out halfway is disastrous.
So model value includes more than peak capability:
- Can one job finish without interruption?
- Can it resume after failure?
- How many jobs finish per week?
- Can the workflow survive a model or vendor change?
4. This is a window for building machinery, not merely consuming AI
Nobody knows whether today's generous AI subscriptions will become more expensive, tighter, cheaper, or simply different.
What we do know is that usage terms are not fixed assets.
Today's unusually favorable access is not a permanent law of nature.
That changes the best use of cheap inference.
The weakest use is to have thousands of useful conversations whose results remain trapped in chat history.
The stronger use is to spend today's inference on systems that reduce tomorrow's inference.
For example:
- automate recurring research and collection;
- turn repeated instructions into explicit contracts and rules;
- replace repeated visual checks with tests and evaluators;
- split giant one-shot jobs into resumable stages;
- store artifacts, evidence, and state outside chat;
- place vendor differences behind a thin adapter;
- record which models actually succeed on which job types.
The principle is simple:
Rent the AI now to build a factory that can keep operating when that exact AI is no longer cheap.
5. Separate disposable usage from durable assets
Ten hours of AI time can leave very different things behind.
| Disposable pattern | Durable pattern |
|---|---|
| Re-explain the same context every time | Store job specifications and rules |
| Ask the model to inspect manually | Build tests and evaluators |
| Run one giant prompt | Add checkpoints and resumable stages |
| Read the answer and move on | Persist artifacts, evidence, and state |
| Use the strongest model for everything | Route cheap first, escalate when needed |
| Grow a vendor-specific magic prompt | Keep a common job contract plus a thin provider adapter |
| Stop when the limit hits | Support retry, resume, and handoff |
Research supports the routing idea. FrugalGPT showed that cascades can select different LLMs per query and, on the evaluated tasks, match the best single model at dramatically lower cost.[4]
A 2026 cost-aware routing study similarly used a cheaper first-stage assignment and escalated only low-quality cases, reporting 97–99% of the strongest model's accuracy while improving serving efficiency.[5]
The strongest model does not need to touch every job.
6. A minimal vendor-portable architecture
You do not need a grand multi-cloud program.
Six separations get most of the benefit.
- Job contract — Define inputs, outputs, and completion criteria without naming a model.
- Provider adapter — Isolate OpenAI-, Anthropic-, or other vendor-specific calls.
- State store — Record where the job stopped.
- Artifact store — Keep code, documents, evidence, and results outside the conversation.
- Evaluator — Test whether "done" is actually done.
- Router — Start with the cheapest capable model and escalate only when evidence says it is necessary.
Microsoft's Well-Architected guidance makes the same general architectural point: reduce tightly coupled dependencies and separate domain logic from infrastructure-specific concerns.[6]
Applied to AI, every sentence of "this only works with Model X" is future migration cost.
7. What should powerful models build while they are cheap?
High-end models are especially valuable when the output keeps producing value after the session ends.
Good targets include:
Automation of recurring work. Research, formatting, testing, publishing, reconciliation, and reporting should not require a fresh human prompt every time.
Recovery mechanisms. Add retries, checkpoints, idempotency, deduplication, and resumability so one quota event does not destroy a long run.
Externalized quality rules. Turn "make it good" into testable criteria instead of relying on one model's taste.
Observability.
Record job type, model, success, retries, duration, and usage.
Model switching. You do not need to run three providers constantly. You do want the door to exist.
Then a future price change becomes a routing adjustment rather than a rewrite.
8. The optimizations that age badly
Cheap abundant AI creates tempting traps.
Do not optimize for "using every last token." Usage is not output.
Do not make giant one-shot jobs the default. They are fragile under quota changes, outages, and model failures.
Do not accumulate excessive vendor-specific prompt folklore. A magical incantation that works only with one model is another form of lock-in.
And do not design around a prediction such as "Vendor A will definitely raise prices" or "Vendor B will always be generous."
You do not need to predict the future if your system can tolerate several futures.
9. Conclusion — subscription generosity is weather; your system is the house
It is reasonable to be annoyed when a $200 plan changes its usage economics.
It is also reasonable to look at another service and think, "I can get far more real work done there."
But if productivity depends entirely on today's pricing table, every vendor announcement becomes an operational incident.
A stronger strategy is to convert today's cheap intelligence into durable capital:
code, tests, evaluators, data, automation, resumable workflows, and replaceable model adapters.
Do not assume today's best model will remain the best.
Do not assume today's cheap subscription will remain cheap.
Build the system while intelligence is cheap enough to build it quickly.
In a world where the "exciting Tuesday announcement" can begin with a long post about tighter usage accounting, portability is not pessimism.
It is simply good engineering.
Sources
- Tibo (@thsottiaux), X post, 2026-09-29, announcing the 2026-09-30 reopening of Pro $200 and the new usage calculation x.com
- OpenAI, “Introducing GPT-6 Sol and Luna”, 2026-09-22 openai.com
- OpenAI API, “Pricing”, checked 2026-09-29 developers.openai.com
- Chen, Lingjiao; Zaharia, Matei; Zou, James, “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance”, Transactions on Machine Learning Research, 2024 openreview.net
- Moslem, Yasmin et al., “Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving”, arXiv, 2026 arxiv.org
- Microsoft Azure Well-Architected Framework, guidance on reducing tightly coupled dependencies and separating domain logic from infrastructure concerns, checked 2026-09-29 learn.microsoft.com
