The easiest mistake in AI coding is optimizing for the price of a single run.
A cheap model that wanders through five dead ends and finishes with “permission denied” can cost more than an expensive model that identifies the real failure mode on the first pass.
In September 2026, Codex gained GPT-6 Sol and GPT-6 Luna alongside GPT-6 Astra. The useful question is no longer “Which model is strongest?”
It is:
Which model minimizes the cost of getting a task all the way to done?
1. What actually launched
OpenAI released GPT-6 Sol and GPT-6 Luna in the API on September 22, 2026, and added them to ChatGPT Work and Codex.[1][2]
GPT-6 Sol is positioned for complex coding and agentic workflows. GPT-6 Luna is optimized for focused, high-volume tasks. Both expose a 1.05M-token context window and up to 128K output tokens.[3][4][5]
GPT-6 Astra sits above them as OpenAI’s model for the hardest end-to-end work, including coding, computer use, research, and document creation.[6]
A practical mental model is:
- Luna: inexpensive production worker
- Sol: default senior operator
- Astra: the person who walks into the server room and immediately asks why that cable is backwards
2. The “GPT-6 came to ChatGPT” naming trap
The September 22 ChatGPT release notes specifically say that GPT-6 Sol and Luna arrived in ChatGPT Work and Codex, and that these models are separate from the models available in Chat.[1]
So these are not equivalent statements:
“GPT-6 Sol is in the ChatGPT product.”
“GPT-6 Sol is selectable in an ordinary Chat conversation.”
The first can be true while the second is false.
Astra is different: its September launch announcement described a rollout to paid ChatGPT users more broadly.[6] The surfaces are different even though all three models share the GPT-6 name.
3. Sticker prices are dramatically different
As of September 23, 2026, the Standard API prices below are per 1M input/output tokens.[3][4][6][7][8][9]
| Model | Input | Output |
|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 |
| GPT-6 Sol | $2.00 | $10.00 |
| GPT-6 Astra | $10.00 | $50.00 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| GPT-5.6 Terra | $2.00 | $12.00 |
| GPT-5.6 Sol | $4.00 | $20.00 |
At identical token counts, Sol is roughly 20× Luna, and Astra is 5× Sol.
That makes “use Luna for everything” look brilliant—until the task requires judgment instead of throughput.
It is the software equivalent of saving on moving costs by carrying a refrigerator across town yourself.
4. A simple 100K-in / 20K-out simulation
Assume one task consumes 100,000 input tokens and 20,000 output tokens. Ignore caching, tool fees, long-context surcharges, and retries.
| Model | Approx. API cost per run |
|---|---|
| GPT-6 Luna | $0.020 |
| GPT-6 Sol | $0.400 |
| GPT-6 Astra | $2.000 |
| GPT-5.6 Luna | $0.044 |
| GPT-5.6 Terra | $0.440 |
| GPT-5.6 Sol | $0.800 |
On raw tokens, Luna wins by a ridiculous margin.
But the more useful formula is:
completion cost ≈ cost per attempt × attempts required + human intervention + lost context
At equal token usage, one Astra run costs about the same as five Sol runs. If a hard task takes Sol five attempts but Astra one, the raw API cost is already even.
Add repeated log reading, re-explaining context, and failed patches, and the break-even point can arrive earlier.
5. The expensive model really can be cheaper
OpenAI’s own Terminal-Bench 4.0 results illustrate this.
GPT-6 Astra scored 57.9% versus 37.3% for GPT-5.6 Sol, while OpenAI reports about 9% lower estimated API cost per task for Astra in the compared settings.[6]
On AutomationBench the comparison was 41.4% versus 18.1%; on internal database-migration tasks, 63.9% versus 42.7%.[6]
So “higher token price” and “higher cost to finish the job” are not the same thing.
Astra also introduces a Codex mechanism for preserving and retrieving context across long sessions, reducing dependence on repeatedly compressing everything into summaries.[6] That matters in debugging and large refactors, where forgetting why the previous fix failed is an excellent way to invent a recurring subscription to the same mistake.
6. Route by task shape, not by cheapest-first escalation
A better strategy is not to force every job through Luna, then Sol, then Astra.
Classify the task first, then escalate after one meaningful failure.
GPT-6 Luna
Use it for:
- grep, extraction, classification, formatting
- tightly specified small edits
- independent batch jobs
- test execution and result summarization
- tasks with mechanical acceptance criteria
GPT-6 Sol
Use it for:
- ordinary multi-file coding
- debugging where the likely area is known
- standard agentic workflows
- research → patch → test loops
- the default Codex workload
GPT-6 Astra
Use it for:
- unknown root cause
- workflows crossing multiple services, permissions, or state systems
- long debugging sessions
- migrations and large refactors
- problems that repeatedly reach 80% and then die
- a task Sol investigated properly once and still could not close
The goal is to avoid “AI overtime”: more exploration, more tokens, more logs, and no additional truth.
7. What happens to GPT-5.6 Terra?
Terra now sits in an awkward price position.
Under Standard pricing:
- GPT-5.6 Terra: $2 input / $12 output
- GPT-6 Sol: $2 input / $10 output
So input pricing is equal while Sol’s output is cheaper.[3][8]
Sol is also explicitly positioned as the newer model for complex coding and agentic workflows.[3]
That does not make Terra useless. A validated production prompt may behave more predictably on a known model, and migration risk is real.
But for a new workload, the clean three-tier map is increasingly:
Luna for volume, Sol for default work, Astra for hard end-to-end work.
8. The real answer comes from 20 tasks of your own
Benchmarks are maps, not your repository.
Split real work into three to five categories and record roughly 20 tasks per category. Track:
- completion without additional human instructions
- time to green
- retries
- tokens or product credits
- amount of logs a human had to inspect
- regressions introduced
- frequency of eventually switching models anyway
Then optimize for:
total cost per completed task
and
completed work per minute of human attention
That is a more useful metric than price per million tokens.
Conclusion: Sol by default, Luna for mechanical volume, Astra for mazes
The current lineup can be simplified.
GPT-6 Luna: high-volume, explicit, independent work.
GPT-6 Sol: normal Codex work.
GPT-6 Astra: ambiguous, long, cross-system problems.
The expensive mistake is often not choosing Astra.
It is spending multiple failed runs proving that you should have chosen Astra.
Token cost matters. So does the time spent reading the same log for the fourth time.
References (9)
- OpenAI, ChatGPT Release Notes, Sep. 22, 2026 help.openai.com
- OpenAI API Changelog, Sep. 22, 2026 developers.openai.com
- OpenAI, GPT-6 Sol model developers.openai.com
- OpenAI, GPT-6 Luna model developers.openai.com
- OpenAI, Models catalog developers.openai.com
- OpenAI, GPT-6 Astra openai.com
- OpenAI, GPT-5.6 Sol model developers.openai.com
- OpenAI, GPT-5.6 Terra model developers.openai.com
- OpenAI, GPT-5.6 Luna model developers.openai.com


