🔍 Read the full analysis: OpenAI’s Half-Price GPT‑6 Sol And Luna Models Keep Benchmark Scores Unchanged on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI has introduced new GPT-6 Sol and Luna models priced at half their previous versions, maintaining similar benchmark scores. While costs are significantly reduced, some evaluation metrics show regressions, especially in knowledge quality.
OpenAI has launched GPT-6 Sol and GPT-6 Luna models on September 22, 2026, with prices halved compared to their GPT‑5.6 predecessors. The models are designed to make advanced AI more accessible for cost-sensitive applications without sacrificing benchmark performance, marking a significant shift in AI deployment economics.
The new GPT‑6 Sol and Luna models are priced at $2.00 and $0.10 per 1 million tokens for input, and $10.00 and $0.50 for output, respectively—roughly 50% less than GPT‑5.6. OpenAI attributes the cost reductions to improvements in caching and inference techniques, which allow these models to be served at lower expense while passing savings onto users.
Independent analysis by Artificial Analysis confirms that, despite the lower costs, the models’ benchmark scores remain stable overall. GPT‑6 Sol scores 48 on the Artificial Analysis Intelligence Index, significantly above the median of 25 for comparable models, with Luna scoring 37 against a median of 12. However, some evaluations show regressions in knowledge tasks, with Sol dropping around 100 Elo points and Luna about 75 in certain benchmarks. These regressions are linked to reduced presentation quality and omitted details, as noted in the analysis.
Quality improvements are evident in hallucination reduction, with Sol decreasing its hallucination rate from 92% to 60%, and Luna from 93% to 77%. However, Sol’s attempt rate also declined from 99% to 83%, leading to fewer answers but fewer errors. The models also demonstrate varied performance in coding tasks, with Sol improving slightly and Luna declining, indicating mixed results across different use cases.
GPT‑6 Sol and Luna: half the price, about the same intelligence
OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.
Per 1M input / output tokens. Cached input reads keep the 90% discount.
Cost per task, halved
Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.
The effort dial moves cost more than the model choice
| Model and effort | Intelligence Index | Cost per task |
|---|---|---|
| GPT‑6 Sol (max) | 48 | $1.06 |
| GPT‑6 Sol (low) | 34 | $0.13 |
| GPT‑6 Luna (max) | 37 | $0.07 |
| GPT‑6 Luna (low) | 21 | $0.0045 |
| GPT‑6 Luna (non‑reasoning) | 18 | $0.01 |
Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.
What got better, and what got worse
Better
- Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
- Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
- OpenAI reports about half as many factual mistakes for Sol as its predecessor
- Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing
Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.
Worse
- GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
- AA‑Briefcase v1.1: Luna down ~45 Elo
- Coding Agent Index: Luna 41, down 2 points
- Both models write more output tokens per task than their predecessors
Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.
What to do about it
Implications for Cost-Effective AI Deployment
The release of GPT‑6 Sol and Luna at half the previous cost could significantly lower barriers for integrating advanced AI into products and workflows. Organizations can now access high-performance models without proportionally increasing expenses, potentially expanding AI adoption across industries. However, the observed regressions in knowledge quality and presentation suggest that cost savings may come with trade-offs in output completeness and accuracy, requiring careful testing for specific applications.
As an affiliate, we earn on qualifying purchases.
Background on OpenAI’s Model Pricing and Performance
OpenAI’s earlier models, including GPT‑5.6, set a high benchmark for AI performance but at a premium cost, limiting widespread deployment in budget-sensitive environments. The company’s announcement of Astra, its top-tier model, emphasized cutting-edge capabilities at a higher price point. The new GPT‑6 Sol and Luna models aim to democratize access by offering similar benchmark scores at a fraction of the cost, leveraging improvements in caching and inference efficiency. Prior to this, reductions in model costs typically coincided with performance drops; however, OpenAI claims these models maintain their benchmark standings despite the price cut.
Independent evaluations have confirmed that while costs have halved, some quality metrics, especially in knowledge accuracy and detailed output, show signs of regression. This indicates a strategic shift toward balancing cost efficiency with acceptable performance levels, rather than pushing for continuous improvements in all metrics.
cost-effective AI chatbot development tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Impact on Real-World Applications
While benchmark scores remain stable, it is still unclear how these models perform in practical, production environments—particularly in tasks requiring high factual accuracy and detailed outputs. The observed regressions in knowledge work suggest that some workflows may need testing before full adoption, especially those relying on comprehensive, well-structured responses.
As an affiliate, we earn on qualifying purchases.
Next Steps and Future Developments
OpenAI is expected to continue refining these models, potentially addressing the identified regressions in knowledge quality. Further independent testing will clarify how well the models perform across varied real-world tasks. Additionally, OpenAI may release updated tools for caching and effort management to optimize deployment further. Industry watchers will monitor whether these cost reductions lead to broader adoption and how users adapt to the trade-offs in output quality.
As an affiliate, we earn on qualifying purchases.
Key Questions
What are GPT‑6 Sol and Luna models?
They are new AI models launched by OpenAI at half the price of previous versions, designed to provide high benchmark performance with lower operational costs.
Do these models perform worse than GPT‑5.6?
Benchmark scores are maintained, but some evaluations, especially in knowledge accuracy and presentation quality, show regressions, indicating mixed performance depending on the task.
How do the cost reductions impact AI deployment?
The significant cost savings could make advanced AI more accessible for a broader range of applications and organizations, potentially expanding AI use cases but requiring careful testing for quality in specific workflows.
Are there trade-offs in using these models?
Yes, while benchmark scores are stable, some quality metrics have declined, particularly in detailed knowledge presentation and hallucination rates, which may affect certain use cases.
What is the significance of caching improvements?
Enhanced caching reduces per-task costs by reusing context more efficiently, enabling lower prices and faster response times, especially for high-volume applications.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
