GPT-6 Sol and Luna Cut Prices in Half, and Coding Scores Fall From 72.7% to 68.8%
OpenAI shipped GPT-6 Sol and Luna on September 22. Input runs $2 per million tokens on Sol and $0.10 on Luna, exactly half the predecessor rate, while the top DeepSWE score drops from 72.7% to 68.8%.
- GPT-6 Sol lists at $2 per million input tokens, half what GPT-5.6 Sol charged.
- Its best DeepSWE score fell from the predecessor's 72.7% to 68.8%.
- Cross
272,000input tokens and the long-context rate erases the discount.
OpenAI put gpt-6-sol and gpt-6-luna on the API on September 22. With GPT-6 Astra, which arrived three weeks earlier, the GPT-6 line now covers three price tiers.
The launch post leads with cost, not capability. "Improvements in caching and inference let us serve these models for less, and we are passing that saving straight through to users," it says. GPT-5.6 Sol charged $4 per million input tokens and $20 per million output. GPT-6 Sol charges $2 and $10. GPT-5.6 Luna charged $0.20 and $1.20; GPT-6 Luna charges $0.10 and $0.50. An OpenAI spokesperson confirmed to VentureBeat that these are permanent list rates, not promotional ones. Claude Opus 5.5, which Anthropic shipped the same day, charges $4 per million input tokens.
| Model (per 1M tokens) | Input | Cached input | Output |
|---|---|---|---|
| GPT-6 Astra | $10 | $1 | $50 |
| GPT-6 Sol | $2 | $0.20 | $10 |
| GPT-5.6 Sol (predecessor) | $4 | $0.40 | $20 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
| GPT-5.6 Luna (predecessor) | $0.20 | $0.02 | $1.20 |
What improved over the predecessor
The launch post names four gains against GPT-5.6.
Factual accuracy moved the most. OpenAI built an internal eval out of real conversations users had flagged as wrong. On that eval, GPT-6 Sol produced "roughly half" the errors its predecessor did. Raise Luna's effort setting, OpenAI says, and it reaches GPT-5.6 Sol's level at one hundredth the cost. The independent lab Artificial Analysis measured the hallucination rate falling from 92% on GPT-5.6 Sol to 60% on GPT-6 Sol.
FrontierCode, which measures whether generated code is mergeable as written, went up. The launch post only calls the gain "substantial" and does not print the two models' scores side by side.
On workflow automation, Luna beats its predecessor by 5.4 percentage points. AutomationBench 1.0.6 checks whether a model can finish sales, marketing, and finance tasks while moving across 47 tools. There, GPT-6 Luna at high effort leads GPT-5.6 Luna by 5.4 points while costing 58% less per task.
The cache discount got deeper. An agent that resends the same context pays less for the portion already processed. GPT-6 applies a 90% discount to those cached input tokens and improves the default hit rate. GitHub reported to OpenAI that over recent months the share of prompt tokens needing fresh processing dropped by more than 50%, which made Copilot responses faster.
The top scores land below the predecessor's
The launch post does not print GPT-5.6 scores for DeepSWE or OSWorld. Every comparison there is against Claude. GPT-6 Sol's DeepSWE v1.1 result of 68.8% appears inside the claim that it comes "within 1.1 points of Claude Fable 5's best score of 69.9% at roughly 80% lower cost per task."
Put the predecessor next to it and the picture changes. GPT-5.6 Sol's best published DeepSWE v1.1 score at its own launch was 72.7%. OSWorld 2.0 also fell, from 65.7% to 60.5%. Those predecessor figures come not from OpenAI but from third-party tabulations (orcarouter, digitalapplied) that placed both models in one table.
| Benchmark | GPT-6 Sol | GPT-5.6 Sol | GPT-6 Luna |
|---|---|---|---|
| DeepSWE v1.1 (real repository tasks) | 68.8% | 72.7% | 66.6% |
| OSWorld 2.0 (computer control) | 60.5% | 65.7% | 58.1% |
| Cost per task (DeepSWE) | $2.74 | $6.46 | Not published |
| Hallucination rate (Artificial Analysis) | 60% | 92% | 77% |
Hold the budget fixed and the ranking flips. Run GPT-6 Sol at high effort and DeepSWE returns 65.3% for $0.64 per task. Run GPT-5.6 Sol at medium and you get 61.1% for $1.42. At equal spend the new model is both cheaper and better. The predecessor only keeps its roughly 4-point lead on work that requires pushing effort to maximum to extract a peak score.
Document work went the other way. Artificial Analysis scored GDPval-AA v2.1, which rates spreadsheet, memo, and slide quality, and found Sol down about 100 Elo and Luna down about 75 Elo, attributing the drop to incomplete deliverables and omitted content. Luna's coding agent index also slipped from 43 to 41, and the output tokens it spends on the same problem rose from 41,000 to 51,000.
Can you use it today
Neither model has a waitlist or an approval step. Access is gated by ChatGPT plan, not by country, and the rates above are USD list prices with no regional tier, so a team in Singapore or Sydney pays what a team in San Francisco pays. What the launch post does not address is where inference runs: OpenAI operates a separate data-residency program for Asia covering ChatGPT Enterprise, Edu, and the API platform, and that program's country list, not this launch, is what answers a processing-location requirement under Singapore's PDPA or the EU AI Act.
| Route | Who gets it | Conditions |
|---|---|---|
| API | Every developer | An API key is all you need |
| ChatGPT Work, Codex | Plus, Pro, Business, Enterprise, Edu | Staged rollout. Check again later if it is missing |
| Desktop app | Free, Go | Luna only |
| Standard chat surface | Not applicable | Not shipped there yet |
The context window is 1.05M tokens with a 128,000-token output ceiling, matching Astra. Knowledge cutoffs are April 20, 2026 for Sol and May 18, 2026 for Luna. Reasoning effort takes six values from none through max, and the default is medium.
One number to check before migrating
The cut does not apply to every request. Once input passes 272,000 tokens, billing switches to the long-context rate and Sol becomes $4 for input and $15 for output. The input rate matches GPT-5.6 Sol at that point, so the halving disappears. A retrieval workload that stuffs whole documents into the prompt collects none of the discount.
Speed splits in two directions as well. Artificial Analysis measured output throughput rising 44%, from 72.6 tokens per second to 104.4, while time to first token stretched into the 100-second range. For a product where a person sits watching the screen until the answer starts streaming, that trade reads as a regression.
What this release opens up is budget, not capability. Artificial Analysis put it as "cost per task halved while the intelligence score stayed at GPT-5.6 levels." At equal spend the new models still win, and the error rate dropped visibly. In the same launch post OpenAI disclosed that its own researchers spend a median of $600 per person per day on tokens, with the 90th percentile at $7,000, priced at API rates. It offered those figures as evidence that a lower price changes the volume of work you can hand a model.
If you already call GPT-5.6 Sol through the API, start by measuring the 95th percentile of input tokens in last week's request logs against 272,000. Below it, changing one model ID halves the bill. Above it, the long-context rate removes most of the reason to move. If the model sits behind a conversational surface, measure time to first response alongside cost before you decide.