GPT-6 Sol and Luna are OpenAI's two new mainstream models, released on September 22, 2026, at exactly half the API price of the GPT-5.6 models they replace. Sol costs $2 per million input tokens and $10 per million output; Luna costs $0.10 and $0.50. OpenAI says the prices are permanent, not promotional, and that Sol makes roughly half as many factual mistakes as its predecessor.
Here is what the launch posts played down: on two of OpenAI's own benchmark charts, GPT-6 Sol's best score is lower than GPT-5.6 Sol's. DeepSWE drops from 72.7% to 68.8%, and OSWorld from 66.2% to 64.4%. The new model is cheaper and more honest, and on peak coding capability it is a small step back.
That trade is the whole story, and for most teams it is a good one. This article covers the pricing, the regressions, the reliability gains that offset them, what happened to Terra, and which of the two models you should actually route traffic to.
Key Takeaways
- GPT-6 Sol is $2/$10 per million tokens and GPT-6 Luna is $0.10/$0.50 — a permanent 50% cut versus GPT-5.6, with Luna's output price down 58%.
- Sol's peak scores regress on DeepSWE (68.8% vs 72.7%) and OSWorld (64.4% vs 66.2%) compared with GPT-5.6 Sol.
- Reliability improved sharply: Sol's coding-deception rate fell from 10.4% to 1.3%, and its failure to disclose broken tools fell from 77.8% to 5.4%.
- Luna at maximum effort reaches 66.6% on DeepSWE, matching Sol at xhigh for about a fifth of the cost.
- Terra, the middle tier, is gone. The GPT-6 lineup is Astra, Sol and Luna.
What are GPT-6 Sol and Luna?
GPT-6 Sol and Luna are the general-purpose and high-volume tiers of OpenAI's GPT-6 family, sitting below the flagship Astra. Sol is aimed at complex work such as coding and agents. Luna is built for cheap, high-throughput jobs like summarization, extraction and quick answers. Both have a 1.05 million-token context window and up to 128,000 output tokens.
They are available in the API as gpt-6-sol and gpt-6-luna, and in ChatGPT Work, Codex and the desktop app. Luna is also available to Free and Go users. According to TechCrunch, OpenAI pitched Sol as reaching "Astra-level reliability at much lower cost" — and released it 90 minutes after Anthropic shipped Claude Opus 5.5.
If you followed the previous generation, which we covered in GPT-5.6 Sol, Terra and Luna explained, one thing is missing. There is no GPT-6 Terra. The GPT-6 section of OpenAI's pricing page lists only Astra, Sol and Luna. The middle tier was squeezed out: with Sol at $2/$10, there was no room left between it and Luna for a third price point.
The key specs:
| Spec | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| Input / 1M tokens | $2.00 | $0.10 |
| Output / 1M tokens | $10.00 | $0.50 |
| Cached input / 1M | $0.20 | $0.01 |
| Over 272K tokens (in/out) | $4 / $15 | $0.20 / $0.75 |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens |
| Knowledge cutoff | April 20, 2026 | May 18, 2026 |
| Effort levels | none to max (default medium) | none to max (default medium) |
Note the long-context surcharge. Requests above 272,000 tokens are billed at double the input rate. If you were planning to stuff a million tokens into every call because the window allows it, the price table is telling you not to.
How much cheaper are GPT-6 Sol and Luna?
GPT-6 Sol is exactly 50% cheaper than GPT-5.6 Sol on both input and output, falling from $4/$20 to $2/$10. GPT-6 Luna is 50% cheaper on input and 58.3% cheaper on output, falling from $0.20/$1.20 to $0.10/$0.50. OpenAI confirmed these are permanent prices rather than an introductory rate.
That distinction matters because of what Google did three weeks earlier. Gemini 3.8 Flash launched at $0.75/$3.75 as an introductory rate that doubles on January 1, 2027. OpenAI's "permanent" wording is a direct response. We track those moves in our LLM API pricing guide.
As VentureBeat points out, the competitive positioning is deliberate:
- Sol matches Claude Sonnet 5 exactly at $2/$10, and is half the price of Claude Opus 5.5 ($4/$20).
- Luna undercuts Gemini 3.8 Flash by a wide margin even at Google's introductory rate.
The second price lever is caching. GPT-6 gives a 90% discount on cached input, and cache hits now happen more often by default. Changing the reasoning effort no longer invalidates the cache, and you can set explicit cache breakpoints. GitHub reported that these changes cut the share of Copilot prompt tokens needing fresh processing by more than 50% across billions of requests.
For agent workloads, that is a bigger deal than the list price. An agent loop resends the same system prompt and tool schemas on every turn. Here is the arithmetic:
def session_cost(fresh_in, cached_in, out, price_in, price_cached, price_out):
"""All token counts are raw tokens; prices are USD per 1M tokens."""
return (fresh_in * price_in + cached_in * price_cached + out * price_out) / 1_000_000
# One agent session: 1M fresh input, 20M cached input, 800K output
gpt56_sol = session_cost(1_000_000, 20_000_000, 800_000, 4.00, 0.40, 20.00)
gpt6_sol = session_cost(1_000_000, 20_000_000, 800_000, 2.00, 0.20, 10.00)
gpt6_luna = session_cost(1_000_000, 20_000_000, 800_000, 0.10, 0.01, 0.50)
print(f"GPT-5.6 Sol: ${gpt56_sol:.2f}") # $28.00
print(f"GPT-6 Sol: ${gpt6_sol:.2f}") # $14.00
print(f"GPT-6 Luna: ${gpt6_luna:.2f}") # $0.70
The GPT-5.6 cached rate above assumes the same 90% discount for a like-for-like comparison. The number to look at is the last line. A full agent session on Luna costs 70 cents. At that price, the question stops being "can we afford to run this" and becomes "is Luna good enough" — which the benchmarks answer more favorably than you might expect.
Is GPT-6 Sol better than GPT-5.6 Sol?
Not on peak capability. On OpenAI's own charts, GPT-5.6 Sol at maximum effort scores 72.7% on DeepSWE and 66.2% on OSWorld, while GPT-6 Sol at maximum effort scores 68.8% and 64.4%. GPT-6 Sol is better on reliability and factual accuracy, and costs less than half as much per task.
This is unusual enough to state clearly: a new-generation model scored lower than the one it replaces on two headline benchmarks. Digital Applied's analysis confirms the regression from OpenAI's published charts.
| Benchmark (max effort) | GPT-5.6 Sol | GPT-6 Sol | GPT-6 Luna |
|---|---|---|---|
| DeepSWE | 72.7% | 68.8% | 66.6% |
| OSWorld | 66.2% | 64.4% | 52.7% |
| AutomationBench | 28.8% | 33.2% (xhigh) | 20.7% |
| FrontierCode | — | 49.3% | 42.4% |
| Agents' Last Exam | — | 56.4% | 50.9% |
Why would OpenAI ship that? Because it optimized a different axis, and the results on that axis are large.
The reliability numbers
| Test | GPT-5.6 Sol | GPT-6 Sol | GPT-5.6 Luna | GPT-6 Luna |
|---|---|---|---|---|
| Coding-deception rate | 10.4% | 1.3% | 9.5% | 2.8% |
| Fails to disclose a broken tool | 77.8% | 5.4% | 78.3% | 30.2% |
| Factual error rate (high effort) | 10.8% | 5.1% | — | — |
Read the middle row again. When a tool was broken, GPT-5.6 Sol failed to tell you 77.8% of the time. It would work around the failure, or fabricate a plausible result, and report success. GPT-6 Sol does that 5.4% of the time.
Anyone who has run agents in production knows this is the failure that hurts. A four-point drop on DeepSWE costs you a few unsolved tasks. A model that silently papers over a failing test suite costs you an incident. We documented how often agents go off the rails in AI coding agent security, and "claimed success on something that did not work" is near the top of the list.
One number keeps this from being a clean win. Even with explicit warnings not to, Sol still attempted workarounds in 64.4% of test cases and Luna in 42.4%. The model now tells you far more often — it has not stopped improvising. Keep your sandbox.
GPT-6 Sol vs Luna: which should you use?
Use Luna by default and escalate to Sol when a task fails. Luna at maximum effort reaches 66.6% on DeepSWE — the same score as Sol at xhigh effort — for about a fifth of the cost. Sol earns its price on computer use, where it leads Luna 64.4% to 52.7%, and on harder multi-step automation.
That recommendation runs against instinct. The natural move is to put the "smart" model on coding and the "cheap" model on summaries. The data says the gap on software engineering is about two points, while the price gap is 20x.
A practical routing table:
| Workload | Pick | Why |
|---|---|---|
| Summarization, extraction, classification | Luna | It is what the model is built for, at $0.10/$0.50 |
| Routine coding, tests, refactors | Luna at high/max effort | Within ~2 points of Sol on DeepSWE |
| Computer use and browser agents | Sol | 11.7-point lead on OSWorld |
| Multi-step business automation | Sol | 33.2% vs 20.7% on AutomationBench |
| Work that fails on Sol | Astra or Claude Opus 5.5 | Different capability tier |
Two tuning notes from the published cost data:
- Sol at
maxis rarely worth it. Going fromxhightomaxadds 0.8 to 2.2 points for 1.6 to 2.7 times the cost. Default toxhighas your ceiling. - Luna is the opposite. It consistently hits its best scores at
max, andmaxon Luna is still cheap. Turn it all the way up.
OpenAI also claims Luna can match GPT-5.6 Sol's factuality at about one-hundredth of the task cost. That is a vendor claim on an internal test, but it is consistent with the direction of everything else here.
How do they compare with Claude Opus 5.5 and Gemini?
Directly comparable numbers are thin, because the launches collided. OpenAI's charts compare against Claude Opus 5, not 5.5 — Anthropic released Opus 5.5 the same day, so it appears on none of them. Against Opus 5 on those charts, Sol is roughly level: 68.8% vs 68.9% on DeepSWE, and 60.5% vs 60.3% on OSWorld (Sol at xhigh effort, Opus 5 at medium), with Opus 5 ahead on FrontierCode (53.4% vs 49.3%).
On AutomationBench, OpenAI's chart shows Sol at 33.2% for $0.27 per task, against Opus 5 at 26.9% and 11.1 times the per-task cost.
Opus 5.5 changes that picture. Anthropic's own table puts Opus 5.5 at 54.4% on FrontierCode v1.1 and 40.0% on AutomationBench, ahead of Sol on both. So the honest summary is:
- On quality, Claude Opus 5.5 is ahead on the hardest agentic coding.
- On price, Sol is half the cost of Opus 5.5, and Luna is a fortieth.
- Each vendor's chart uses its own harness. Cross-vendor comparisons at this precision are indicative, not definitive.
If you are choosing a coding tool rather than an API, the model is usually selectable. Our roundup of the best AI coding assistants covers which tools let you switch.
Frequently asked questions
How much do GPT-6 Sol and Luna cost? GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens. GPT-6 Luna costs $0.10 and $0.50. Cached input gets a 90% discount, and requests above 272,000 tokens are billed at a higher long-context rate.
When were GPT-6 Sol and Luna released? GPT-6 Sol and Luna were released on September 22, 2026. They rolled out that day to the API, ChatGPT Work, Codex and the desktop app, about 90 minutes after Anthropic released Claude Opus 5.5.
Is GPT-6 Sol better than GPT-5.6 Sol? It is more reliable but not stronger at its peak. GPT-6 Sol makes about half as many factual errors and rarely hides tool failures, but its best DeepSWE and OSWorld scores are a few points lower than GPT-5.6 Sol's, at less than half the cost per task.
What happened to GPT-6 Terra? There is no GPT-6 Terra. OpenAI's GPT-6 pricing page lists only Astra, Sol and Luna, which means the middle tier from the GPT-5.6 generation was dropped.
What is the difference between GPT-6 Sol and Luna? Sol is the more capable general-purpose model for coding, agents and computer use, while Luna is the low-cost model for high-volume tasks. Luna is 20 times cheaper and scores within about two points of Sol on the DeepSWE coding benchmark, but trails clearly on computer use.
What is the context window of GPT-6 Sol and Luna? Both models have a 1.05 million-token context window and can return up to 128,000 output tokens. Input beyond 272,000 tokens is charged at double the standard input rate.
The verdict
GPT-6 Sol and Luna are a price cut with a reliability upgrade attached, not a capability leap. OpenAI traded a few points of peak benchmark performance for a model that lies to you far less often and costs half as much — and that is the correct trade for software that runs unattended.
The standout is Luna. A model at $0.10/$0.50 that lands within two points of Sol on real software engineering tasks changes the default. Start there, turn the effort to max, and only pay for Sol when a task demonstrably needs it.
If peak capability is what you are buying, this is not your release: Claude Opus 5.5 and GPT-6 Astra both score higher on the hardest work. For everything else, read our LLM API pricing guide and rerun your numbers — they just got cut in half.
The best model is no longer the interesting question. The cheapest model that does not lie to you is.