GPT-6 Sol vs Luna: Which OpenAI Budget Model Should You Use?
We compared GPT-6 Sol and GPT-6 Luna the practical way: same prompts, three models, live timers and running costs, following AI with Surya's full side-by-side arena session, with official pricing charts from Matthew Berman's launch walkthrough cross-checked against OpenAI's announcement. Sol is the mid-priced step below flagship GPT-6 Astra, while Luna, at $0.10 per million input tokens, is the budget workhorse. Below you will see where each model lives, what they cost in practice, where Sol pulls ahead of Luna, and when Astra is still worth paying for.
Source & credits
Screenshots in this guide are captured from AI with Surya's public walkthrough video. Every step links back to the exact moment it shows, so you can follow along.
AI with Surya ↗What shipped on September 22
- 1
Know the lineup: three models launched the same day
September 22 was a three-model day: Anthropic shipped Claude Opus 5.5, and OpenAI answered with GPT-6 Sol and GPT-6 Luna within hours. Both new OpenAI models sit under flagship GPT-6 Astra, which arrived earlier in September, and both were trained with similar methods. The pitch is simple: keep most of Astra's capability, cut the price dramatically.

Opus 5.5, Sol and Luna all landed within hours of each other.Watch at 0:03 - 2
See where Sol and Luna land on real business tasks
The launch-day numbers put the three models on one ladder. On the AutomationBench business-tasks test, Claude Opus 5.5 finished 40 percent of tasks at $20 per million output tokens, GPT-6 Sol reached 33 percent at $10, and GPT-6 Luna about 21 percent at $0.50. Sol gives up 7 points to Opus at half the price, and Luna gives up more score still while costing almost nothing per token.

One ladder, three price points: 40%, 33% and about 21% of tasks finished.Watch at 2:07 - 3
Register the price cut: both cost half of GPT-5.6
OpenAI cut API prices by 50 percent across the pair, on top of the 80 percent drop GPT-5.6 Luna had already received. GPT-6 Sol moves from $4 to $2 per million input tokens and from $20 to $10 of output; GPT-6 Luna moves from $0.20 to $0.10 input and $1.20 to $0.50 output. Per token, that makes Sol a mid-tier workhorse and Luna one of the cheapest frontier-family models ever listed.

Sol $2/$10, Luna $0.10/$0.50 per million tokens: both 50% below GPT-5.6.Watch at 1:26
The three-model arena setup
- 4
Set up the same three-way comparison yourself
The comparison session ran in a custom Model Arena page that fires one prompt at Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna simultaneously, with tokens, time and running cost tallied per model. All three models were accessed through OpenRouter, so you can reproduce the exact setup without an OpenAI key. Identical settings across the columns keep the race fair.

One prompt in, three races out: timers and costs tracked per model.Watch at 2:30 - 5
Write the prompt once, run all three
Each task is typed into the prompt builder once and dispatched to every model at once, first try, no retries. The first brief asks for a one-page site for a fictional umbrella brand called Squall, with animation, customer reviews and a pre-order section, using no external images. Watching the three columns stream side by side is the fastest way to feel the speed gap between Luna, Sol and Opus.

The Squall brief: one page, animation, reviews, pre-order, zero external images.Watch at 6:45
Price and speed in practice
- 6
Read test one: a landing page for pennies or dollars
The umbrella landing page already tells the pricing story. GPT-6 Luna finished in 1 minute 22 seconds for $0.0005, GPT-6 Sol took 1 minute 48 seconds for $0.18, and Claude Opus 5.5 needed 3 minutes 25 seconds at $0.83. All three shipped a working page with reviews and a pre-order block; the difference was polish, with Opus adding the richest storm animation and Sol adding audio feedback.

Same page, three bills: $0.0005, $0.18 and $0.83.Watch at 6:00 - 7
Read test two: the storm dashboard cost gap
The storm-tracking dashboard widens the gap. GPT-6 Luna delivered in 1 minute 41 seconds for $0.0073, GPT-6 Sol in 1 minute 49 seconds for $0.12, and Claude Opus 5.5 spent 9 minutes 17 seconds, burned 30.0k thinking plus 63.7k output tokens, and billed $1.27. That is more than 150 times the cost of Luna for the same brief and about ten times Sol, in exchange for the most polished result of the three.

Opus: 9m17s and $1.27. Sol: $0.12. Luna: $0.0073.Watch at 10:35 - 8
Check the official API prices before you commit
OpenAI's own pricing table is the number to plan against: gpt-6-sol runs $2 per million input tokens and $10 of output, while gpt-6-luna lists $0.10 input and $0.50 output. Both lines carry a 50 percent cheaper badge against their GPT-5.6 predecessors. The same announcement keeps GPT-6 Astra at the top of the lineup as the model to choose when you want the best results and an uncompromising experience.

Model IDs: gpt-6-sol and gpt-6-luna. Astra stays the uncompromising option.Watch at 0:45 - 9
Put all three side by side, including Opus
Charting the September 22 lineup against Claude Opus 5.5 makes the tiers obvious. Input price per million tokens runs Luna $0.10, Sol $2.00 and Opus 5.5 $4.00, and output runs $0.50, $10.00 and $20.00. At max effort the published benchmarks keep the same ordering: FrontierCode v1.1 Main scores 42.4, 49.3 and 54.4 percent, and AutomationBench scores 20.7, 33.2 and 40.0 percent.

Price tiers and score tiers move together: Luna, Sol, Opus 5.5.Watch at 5:02
Quality where it counts: one dashboard, one 3D game
- 10
Watch Sol reroute trucks while Luna gets stuck
Quality separates the two budget models long before benchmarks do. On the storm dashboard, Luna's version kept routing trucks back into the weather and never recovered the schedule, while Sol marked every truck inside the storm red and its redirect control actually re-plotted routes, cutting at-risk deliveries. For any workflow where the model has to reason about state, this is the gap that matters.

Sol's dashboard: trucks in the storm turn red and rerouting works.Watch at 8:15 - 11
Grade test three: a 3D game in one prompt
Storm Run, a self-contained 3D sailing game, was the hardest single prompt. GPT-6 Luna produced something playable in 1 minute 30 seconds for $0.0075, GPT-6 Sol spent 2 minutes 17 seconds and $0.18, and Claude Opus 5.5 took 18 minutes 7 seconds, 74.9k thinking plus 123.1k output tokens, and $2.47. Cost and ambition scale together, and the screenshots below show whether the extra money bought anything.

One 3D game, three budgets: $0.0075, $0.18, $2.47.Watch at 12:15 - 12
See Luna's ceiling: fast, cheap, barebones
Luna's Storm Run loads instantly and the HUD works, but the ocean is a flat gray plane with no water texture, rain streaks are the only weather, and there is little to actually play. It is the same pattern as the dashboard test: Luna ships the shape of the task and stops. For prototypes, drafts and internal tools, that is often all you need.

Luna's boat sails at 8 km/h on an ocean that is barely there.Watch at 11:58 - 13
See Sol's step up: a game you can actually play
Sol's build has real water, wave motion, a wind system and a course: eight buoys to round, an orange glow marking the next one, and a bearing readout counting down the meters. It is the difference between a mock-up and something a colleague could play. Sol at $0.18 bought gameplay; Luna at $0.0075 bought a loading screen that moves.

Follow the orange glow: Sol's build is an actual game.Watch at 12:45 - 14
See what the flagship's extra $2.29 buys
Opus 5.5's Storm Run is the showcase: dynamic ocean water, sunlight scattering through storm clouds, a camera that surges with the waves, and a killer-wave event the game warns you about. None of that exists in the two cheaper builds. It is the clearest picture of what the expensive tier is for: work where quality is the product.

Killer wave ahead: the flagship tier renders weather the budget tiers cannot.Watch at 14:15
GPT-6 Astra vs Luna: when the flagship still wins
- 15
Read the Astra line before downgrading everything
On OpenAI's FrontierCode chart, Astra's curve sits above Sol's at every effort level, and even Astra at low effort scores 45.3 percent at $1.70 per task. Compare that with Luna at max effort, 42.4 percent at about $0.11 per task, and the rule appears: Luna trades a few points of capability for roughly 15 times lower cost. When a task sits near that line, Astra's low-effort setting is a middle path worth testing.

Astra low effort: 45.3% at $1.70 per task, the middle path.Watch at 4:16 - 16
Decide with the Astra vs Luna matrix
Computer use is where the Astra vs Luna gap is widest: on OSWorld 2.0 offline tasks, Luna at max effort manages 52.7 percent at $0.27 per task while Astra's curve tops out around 73 percent. The working rule from the launch coverage: hand Luna 90 to 95 percent of everyday volume work, give Sol the quality-sensitive automation that balks at Astra's price, and keep Astra for the hardest reasoning, coding and agent work, where a wrong answer costs more than the tokens.

Luna max: 52.7%. Astra: roughly 73%. Computer use is the widest gap.Watch at 4:46

