GPT-6.1 Sol vs Claude Sonnet 5.5: The Same-Price Showdown
Anthropic shipped Claude Sonnet 5.5 on September 28 and OpenAI answered with GPT-6.1 Sol at DevDay on September 29 — and the awkward part is that both cost exactly $2 per million input tokens and $10 output. OpenAI's launch comparisons only named Opus; Anthropic's only named Sonnet 5. So the matchup that actually matters went untested by either lab. This guide assembles it from third-party effort-by-effort scores, both price cards, and same-prompt build tests, then ends with a routing rule you can apply to your own workload today.
Source & credits
Screenshots in this guide are captured from The Intelligence Desk's public walkthrough video. Every step links back to the exact moment it shows, so you can follow along.
The Intelligence Desk ↗Two Mid-Tier Flagships, 24 Hours Apart
- 1
GPT-6.1 Sol arrives at DevDay: near-Astra work at a fifth of Astra's price
OpenAI's announcement page leads with a glowing sun and one promise: nearly GPT-6 Astra's intelligence on agentic coding, computer use and professional work, at one-fifth of Astra's standard token prices. It shipped September 29, one week after GPT-6 Sol, and hours after the Wall Street Journal reported that the pricier GPT-6.1 Astra was scrapped over safety-eval failures. Developers call it as gpt-6.1-sol; it lives in ChatGPT Work and Codex, not yet in regular Chat.

OpenAI's own hero image: a sun named Sol, priced like a workhorse.Watch at 0:00 - 2
Sonnet 5.5 came out the day before — at Sol's exact price
Claude Sonnet 5.5 launched September 28 priced the same as Sonnet 5: $2 per million input tokens, $10 output, $0.20 for cache reads — with Anthropic claiming up to 30% lower cost per task than its predecessor. That puts it level with GPT-6.1 Sol per token and half of Opus 5.5's $4/$20. Anthropic's own announcement says Sonnet 5.5 runs 30% faster and costs up to 30% less for most work. OpenAI's launch charts compared Sol against Opus and never mentioned Sonnet — which is why this page exists.

The collision nobody staged: two $2/$10 models, one day apart.Watch at 4:35 - 3
Where Sol sits in OpenAI's lineup — and what it's aimed at
OpenAI's own price table slots GPT-6.1 Sol between the flagship and the budget tier: $2 in, $0.10 cached input, $10 out per million short-context tokens (up to 272K input), then $4 and $15 in long-context territory. Astra stays at $10/$50; Luna starts at $0.10/$0.50. Sol is aimed at three workloads — agentic coding, computer use and office documents — which is precisely Sonnet 5.5's territory. At debut Sonnet 5.5 scored 56 on the Artificial Analysis Index, three points above GPT-6 Astra's 53, so OpenAI's new workhorse had something to prove.

Flagship, workhorse, budget row — Sol is the middle card with a point to prove.Watch at 1:35
Same Sticker, Very Different Bills
- 4
Same per-token price — but Sonnet ships at weaker default settings
Per million tokens the two are twins: $2 in, $10 out. The scores are not. At Anthropic's shipping defaults — Medium effort in the Claude apps and Claude Code, High on the platform API — the Artificial Analysis snapshot has GPT-6.1 Sol ahead at every one: 48 vs 41 at Medium, 50 vs 47 at High, and even Sol's Low setting (42) beats Sonnet's Low (36). Only when Sonnet is pushed past its defaults, to Xhigh (52) and Max (56), does it climb past Sol's 51 and 52 ceiling.

Twins on price, not on defaults — Sol wins where most people actually run.Watch at 5:30 - 5
The bill: Sol does the same work for about a third of the money
Identical token prices produce wildly different invoices because Sonnet 5.5 is far more talkative: across the full index it wrote 410 million output tokens against Sol's 67 million — roughly 193,000 per task versus 31,000, six times the tokens at the same $10 rate. Per task that lands at $0.21 vs $0.59 at Medium and $0.32 vs $1.08 at High — about a third of the bill for a higher score. Cached input widens it further for agent loops: Sol charges $0.10 per million, Sonnet $0.20.

Same rate card, one-third the bill — verbosity is the whole difference.Watch at 6:05 - 6
Turn both dials to max and Sonnet takes the lead back — at 10x the price
Push both models past their defaults and the picture inverts. At Xhigh, Sonnet edges Sol by a single point (52 vs 51) while its per-task cost jumps to $2.74 against $0.39. At Max, Sonnet wins by four points (56 vs 52) at $7.60 versus $0.72 — roughly ten times the price per task. If you truly need that 56-level score, note the third card on this chart: Opus 5.5 reaches 56 at Xhigh for about $3.50 a task, less than half of Sonnet's Max bill.

Sonnet's comeback is real — and it costs about ten times more per task.Watch at 7:05
The Effort Ladder Decides Everything
- 7
The catch in one picture: 56 vs 52 at Max
Here is the entire matchup compressed into one slide. At the highest setting the Artificial Analysis Index gives Sonnet 5.5 a 56 to Sol's 52 — but a Sonnet task at Max costs $7.60 against Sol's $0.72, about ten times more. Run Sonnet at its normal settings instead, and GPT-6.1 Sol is both smarter and much cheaper. Everything else on this page is detail hanging off that trade: you are buying Sonnet's top end, or Sol's everyday value.

One bar chart is the whole debate: Sonnet's ceiling against Sol's price.Watch at 0:45 - 8
Reading OpenAI's launch numbers honestly: Opus is still the ceiling
Against Anthropic's flagship, Sol doesn't win — Opus 5.5 matches or beats it at every effort level (58/56/54/51/42 vs 52/51/50/48/42). But OpenAI's own launch card claims real scalps: DeepSWE v1.1 at 75.2% on high effort, 1.1 points above GPT-6 Astra; OSWorld 2.0 computer use at 71.4% for $1.27 a task where Astra pays $9.44; GDP.pdf at 32.0% over Opus 5.5's 28.8%; AutomationBench ahead of Opus at medium effort (31.7% vs 29.5%) though behind it at max. Sol also reads a 1.05M-token window with a 128K output cap.

Great deal next to Opus — but the ceiling belongs to Anthropic either way.Watch at 3:30 - 9
To be fair: Sol's best beats Opus's medium at half the cost
The fairest way to read the ladder: Sol at Max scores 52 for $0.72 per task, while Opus at its Medium default scores 51 for $1.34 — a higher score for about half the money. On a science terminal test OpenAI's numbers put Sol at $5.47 per task against Opus 5.5's $23.21 at max effort. The verdict stamps say it plainly: a great deal, not an Opus killer. For this page the takeaway is simpler — Sonnet 5.5 lives in the same middle band, and Sol prices that band aggressively.

Sol's ceiling overlaps Opus's floor at a fraction of the spend.Watch at 4:05
Same Prompt, Both Models: Real Builds
- 10
The lab: same prompts, one harness, both models
Scores get you shortlisted; builds get you signed. AIex's AI Workbench ran both models through one local benchmark harness — a claude-code runner for Sonnet and a codex-cli runner for Sol, each on its own subscription — with LLM-judged functional and visual scores. His console shows how close the two finished: GPT-6.1 Sol weighted 5.7/10, Claude Sonnet 5.5 5.6/10. Note his own caveat on screen: scores are not comparable one-to-one, because each Sol task was sampled five times against one Sonnet run.

A home-grown Arena: one harness, two subscriptions, photo-finish averages.Watch at 1:20 - 11
Build test: Sol's VANTA wind-tunnel configurator
Given the same 3D product-page brief, GPT-6.1 Sol produced this VANTA showroom: a silver roadster on a turntable inside a dark wind tunnel, a turbine wall behind it, and a working spec-instrument panel with drag and downforce readouts ticking along the bottom. The address bar tags the build model=gpt-6.1-sol. It is the moodier of the two takes — cinematic lighting, restrained palette, live instrumentation instead of decoration.

Sol's answer to the same brief: dark, cinematic, instrumented.Watch at 1:50 - 12
Build test: Sonnet's flyable island stunt game
On a flying-game prompt, Claude Sonnet 5.5 delivered a playable side-scroller — a biplane banking over palm-fringed islands with a wind gauge in the cockpit, pick-up tokens drifting in the air and a live telemetry strip along the bottom. The URL tags it model=claude-sonnet-5.5. It runs, it animates, and the physics feel tuned rather than generated — exactly the front-end polish Sonnet's reputation is built on.

Sonnet's take: instantly playable, lovingly tuned, colorful.Watch at 4:00 - 13
Creativity test: Sonnet's carnival acrobat catcher
For the open-ended creativity task Sonnet 5.5 built a carnival shooting-gallery where you catch acrobats on a seesaw: a LEVEL 2 banner, a score of 3,600 with 13 of 16 placed, waiting-acrobat cards weighted heavy and medium, an autopilot demo mode and a physics tilt meter. The URL confirms model=claude-sonnet-5.5. It is game-design-literate — scoring, fail states, flavor text — the kind of composed interactivity Sonnet consistently ships.

Levels, score, autopilot, physics — Sonnet ships a whole game loop.Watch at 6:00 - 14
Creativity test: Sol's Wind/Weft sail studio
The same creative brief sent GPT-6.1 Sol somewhere stranger and more art-directed: Wind/Weft, a sail-design studio where colored thread paths converge into a face before a wind-sovereignty panel with tension readouts and a live mode. The address bar tags model=gpt-6.1-sol. Where Sonnet built a game, Sol built a mood — fewer affordances, stronger identity, the kind of output you show clients to argue models have taste.

Sol's take: an art piece with a settings panel — taste over playability.Watch at 6:25
Which One Should You Pick?
- 15
The routing rule: match the model to the effort setting you actually use
The decision tree writes itself from the data above. Run Sonnet 5.5 at Low, Medium or High — the settings most coding agents, browser-automation and office workflows use — and GPT-6.1 Sol will likely score higher for about a third of the bill, so pilot it this week. If you deliberately push Sonnet to Max, compare against Opus 5.5, which reaches the same 56 at Xhigh for roughly half the per-task cost. And for the hardest, longest jobs, Opus 5.5 remains the ceiling nobody beat in September.

Three arrows, one habit: check your effort setting before you check your loyalty.Watch at 8:58 - 16
Before you migrate: three limits worth knowing
Three caveats temper the Sol hype. It is not in regular ChatGPT chat yet — you reach it through ChatGPT Work, Codex and the API. Its Ultrafast speed tier is not out either: OpenAI's announcement promises up to 8x faster generation in Codex, while VentureBeat reports Astra's live Ultrafast runs at 6x standard rates, so expect a premium. And these index scores are days old — snapshots move. Offsetting note from launch coverage: OpenAI says Sol's error rate stays within 1.9% of GPT-6 Astra with fewer safety-reviewer circumvention attempts than GPT-6 Sol.

Buy the model you can reach today; Ultrafast is a promise, not a product.Watch at 8:30
Frequently asked questions
Keep exploring
- ChatGPT tool profile: plans, models and setup
- Claude tool profile: plans, models and setup
- Claude Sonnet 5.5 vs GPT-6: the older OpenAI lineup tested
- Claude Sonnet 5.5 vs Gemini 3.8 Flash: the budget-model matchup
- How to use GPT-6.1 Sol: the full getting-started guide
- GPT-6.1 Sol vs Gemini 3.8: the other Sol head-to-head

