GPT-6.1 Sol vs Claude Sonnet 5.5: The Same-Price Showdown

Anthropic shipped Claude Sonnet 5.5 on September 28 and OpenAI answered with GPT-6.1 Sol at DevDay on September 29 — and the awkward part is that both cost exactly $2 per million input tokens and $10 output. OpenAI's launch comparisons only named Opus; Anthropic's only named Sonnet 5. So the matchup that actually matters went untested by either lab. This guide assembles it from third-party effort-by-effort scores, both price cards, and same-prompt build tests, then ends with a routing rule you can apply to your own workload today.

Source & credits

Screenshots in this guide are captured from The Intelligence Desk's public walkthrough video. Every step links back to the exact moment it shows, so you can follow along.

The Intelligence Desk ↗

Two Mid-Tier Flagships, 24 Hours Apart

  1. 1

    GPT-6.1 Sol arrives at DevDay: near-Astra work at a fifth of Astra's price

    OpenAI's announcement page leads with a glowing sun and one promise: nearly GPT-6 Astra's intelligence on agentic coding, computer use and professional work, at one-fifth of Astra's standard token prices. It shipped September 29, one week after GPT-6 Sol, and hours after the Wall Street Journal reported that the pricier GPT-6.1 Astra was scrapped over safety-eval failures. Developers call it as gpt-6.1-sol; it lives in ChatGPT Work and Codex, not yet in regular Chat.

    OpenAI announcement page for GPT-6.1 Sol showing a glowing sun hero between the words GPT-6.1 and Sol above the navigation bar.
    OpenAI's own hero image: a sun named Sol, priced like a workhorse.Watch at 0:00
  2. 2

    Sonnet 5.5 came out the day before — at Sol's exact price

    Claude Sonnet 5.5 launched September 28 priced the same as Sonnet 5: $2 per million input tokens, $10 output, $0.20 for cache reads — with Anthropic claiming up to 30% lower cost per task than its predecessor. That puts it level with GPT-6.1 Sol per token and half of Opus 5.5's $4/$20. Anthropic's own announcement says Sonnet 5.5 runs 30% faster and costs up to 30% less for most work. OpenAI's launch charts compared Sol against Opus and never mentioned Sonnet — which is why this page exists.

    Chart slide titled OpenAI picked the wrong fight comparing Claude Opus 5.5 at 4 and 20 dollars per million tokens with GPT-6.1 Sol and Claude Sonnet 5.5 both at 2 and 10 dollars.
    The collision nobody staged: two $2/$10 models, one day apart.Watch at 4:35
  3. 3

    Where Sol sits in OpenAI's lineup — and what it's aimed at

    OpenAI's own price table slots GPT-6.1 Sol between the flagship and the budget tier: $2 in, $0.10 cached input, $10 out per million short-context tokens (up to 272K input), then $4 and $15 in long-context territory. Astra stays at $10/$50; Luna starts at $0.10/$0.50. Sol is aimed at three workloads — agentic coding, computer use and office documents — which is precisely Sonnet 5.5's territory. At debut Sonnet 5.5 scored 56 on the Artificial Analysis Index, three points above GPT-6 Astra's 53, so OpenAI's new workhorse had something to prove.

    Lineup slide listing GPT-6 Astra the flagship at 10 and 50 dollars, GPT-6.1 Sol the workhorse at 2 and 10 dollars and GPT-6 Luna at 0.10 and 0.50 dollars beside OpenAI's official price table.
    Flagship, workhorse, budget row — Sol is the middle card with a point to prove.Watch at 1:35

Same Sticker, Very Different Bills

  1. 4

    Same per-token price — but Sonnet ships at weaker default settings

    Per million tokens the two are twins: $2 in, $10 out. The scores are not. At Anthropic's shipping defaults — Medium effort in the Claude apps and Claude Code, High on the platform API — the Artificial Analysis snapshot has GPT-6.1 Sol ahead at every one: 48 vs 41 at Medium, 50 vs 47 at High, and even Sol's Low setting (42) beats Sonnet's Low (36). Only when Sonnet is pushed past its defaults, to Xhigh (52) and Max (56), does it climb past Sol's 51 and 52 ceiling.

    Score comparison listing GPT-6.1 Sol at 42, 48, 50, 51 and 52 by effort against Claude Sonnet 5.5 at 36, 41, 47, 52 and 56 under a note reading same per token 2 dollars in and 10 dollars out.
    Twins on price, not on defaults — Sol wins where most people actually run.Watch at 5:30
  2. 5

    The bill: Sol does the same work for about a third of the money

    Identical token prices produce wildly different invoices because Sonnet 5.5 is far more talkative: across the full index it wrote 410 million output tokens against Sol's 67 million — roughly 193,000 per task versus 31,000, six times the tokens at the same $10 rate. Per task that lands at $0.21 vs $0.59 at Medium and $0.32 vs $1.08 at High — about a third of the bill for a higher score. Cached input widens it further for agent loops: Sol charges $0.10 per million, Sonnet $0.20.

    Effort table scoring GPT-6.1 Sol at 42, 48 and 50 against Claude Sonnet 5.5 at 36, 41 and 47 with per-task costs of 0.21 against 0.59 dollars and a bar chart showing 410 million output tokens for Sonnet against 67 million for Sol.
    Same rate card, one-third the bill — verbosity is the whole difference.Watch at 6:05
  3. 6

    Turn both dials to max and Sonnet takes the lead back — at 10x the price

    Push both models past their defaults and the picture inverts. At Xhigh, Sonnet edges Sol by a single point (52 vs 51) while its per-task cost jumps to $2.74 against $0.39. At Max, Sonnet wins by four points (56 vs 52) at $7.60 versus $0.72 — roughly ten times the price per task. If you truly need that 56-level score, note the third card on this chart: Opus 5.5 reaches 56 at Xhigh for about $3.50 a task, less than half of Sonnet's Max bill.

    Receipt slide showing GPT-6.1 Sol costs from 0.13 to 0.72 dollars per task beside Claude Sonnet 5.5 costs from 0.41 to 7.60 dollars with the full effort table annotated plus one point at seven times the cost and plus four points at ten times the price.
    Sonnet's comeback is real — and it costs about ten times more per task.Watch at 7:05

The Effort Ladder Decides Everything

  1. 7

    The catch in one picture: 56 vs 52 at Max

    Here is the entire matchup compressed into one slide. At the highest setting the Artificial Analysis Index gives Sonnet 5.5 a 56 to Sol's 52 — but a Sonnet task at Max costs $7.60 against Sol's $0.72, about ten times more. Run Sonnet at its normal settings instead, and GPT-6.1 Sol is both smarter and much cheaper. Everything else on this page is detail hanging off that trade: you are buying Sonnet's top end, or Sol's everyday value.

    Slide titled The catch showing index scores of 56 for Sonnet 5.5 against 52 for GPT-6.1 Sol at max effort with cost per task of 7.60 dollars against 0.72 dollars marked ten times more.
    One bar chart is the whole debate: Sonnet's ceiling against Sol's price.Watch at 0:45
  2. 8

    Reading OpenAI's launch numbers honestly: Opus is still the ceiling

    Against Anthropic's flagship, Sol doesn't win — Opus 5.5 matches or beats it at every effort level (58/56/54/51/42 vs 52/51/50/48/42). But OpenAI's own launch card claims real scalps: DeepSWE v1.1 at 75.2% on high effort, 1.1 points above GPT-6 Astra; OSWorld 2.0 computer use at 71.4% for $1.27 a task where Astra pays $9.44; GDP.pdf at 32.0% over Opus 5.5's 28.8%; AutomationBench ahead of Opus at medium effort (31.7% vs 29.5%) though behind it at max. Sol also reads a 1.05M-token window with a 128K output cap.

    Intelligence Index table showing Opus 5.5 at 58, 56, 54, 51 and 42 by effort above GPT-6.1 Sol at 52, 51, 50, 48 and 42 stamped Opus greater than or equal to Sol every setting.
    Great deal next to Opus — but the ceiling belongs to Anthropic either way.Watch at 3:30
  3. 9

    To be fair: Sol's best beats Opus's medium at half the cost

    The fairest way to read the ladder: Sol at Max scores 52 for $0.72 per task, while Opus at its Medium default scores 51 for $1.34 — a higher score for about half the money. On a science terminal test OpenAI's numbers put Sol at $5.47 per task against Opus 5.5's $23.21 at max effort. The verdict stamps say it plainly: a great deal, not an Opus killer. For this page the takeaway is simpler — Sonnet 5.5 lives in the same middle band, and Sol prices that band aggressively.

    To be fair slide showing Sol's best max score at 0.72 dollars per task against Opus at medium 1.34 dollars and a science terminal test bar of 5.47 dollars against 23.21 dollars labeled great deal not an Opus killer.
    Sol's ceiling overlaps Opus's floor at a fraction of the spend.Watch at 4:05

Same Prompt, Both Models: Real Builds

  1. 10

    The lab: same prompts, one harness, both models

    Scores get you shortlisted; builds get you signed. AIex's AI Workbench ran both models through one local benchmark harness — a claude-code runner for Sonnet and a codex-cli runner for Sol, each on its own subscription — with LLM-judged functional and visual scores. His console shows how close the two finished: GPT-6.1 Sol weighted 5.7/10, Claude Sonnet 5.5 5.6/10. Note his own caveat on screen: scores are not comparable one-to-one, because each Sol task was sampled five times against one Sonnet run.

    Benchmark console showing claude-code and codex-cli model runners with recent runs scoring GPT-6.1 Sol at 5.7 out of 10 and Claude Sonnet 5.5 at 5.6 out of 10 on weighted judge averages.
    A home-grown Arena: one harness, two subscriptions, photo-finish averages.Watch at 1:20
  2. 11

    Build test: Sol's VANTA wind-tunnel configurator

    Given the same 3D product-page brief, GPT-6.1 Sol produced this VANTA showroom: a silver roadster on a turntable inside a dark wind tunnel, a turbine wall behind it, and a working spec-instrument panel with drag and downforce readouts ticking along the bottom. The address bar tags the build model=gpt-6.1-sol. It is the moodier of the two takes — cinematic lighting, restrained palette, live instrumentation instead of decoration.

    VANTA car showroom built by GPT-6.1 Sol showing a silver roadster on a platform inside a wind tunnel with a turbine wall and a spec instrument panel with drag and downforce readouts.
    Sol's answer to the same brief: dark, cinematic, instrumented.Watch at 1:50
  3. 12

    Build test: Sonnet's flyable island stunt game

    On a flying-game prompt, Claude Sonnet 5.5 delivered a playable side-scroller — a biplane banking over palm-fringed islands with a wind gauge in the cockpit, pick-up tokens drifting in the air and a live telemetry strip along the bottom. The URL tags it model=claude-sonnet-5.5. It runs, it animates, and the physics feel tuned rather than generated — exactly the front-end polish Sonnet's reputation is built on.

    Side-scrolling biplane game built by Claude Sonnet 5.5 with the plane banking over a palm tree island, a wind gauge HUD, floating token pickups and a telemetry bar at the bottom.
    Sonnet's take: instantly playable, lovingly tuned, colorful.Watch at 4:00
  4. 13

    Creativity test: Sonnet's carnival acrobat catcher

    For the open-ended creativity task Sonnet 5.5 built a carnival shooting-gallery where you catch acrobats on a seesaw: a LEVEL 2 banner, a score of 3,600 with 13 of 16 placed, waiting-acrobat cards weighted heavy and medium, an autopilot demo mode and a physics tilt meter. The URL confirms model=claude-sonnet-5.5. It is game-design-literate — scoring, fail states, flavor text — the kind of composed interactivity Sonnet consistently ships.

    Carnival acrobat game built by Claude Sonnet 5.5 showing a striped big-top tent, a level 2 score of 3,600 with 13 of 16 acrobats placed, waiting acrobat cards and a tilt and balance meter.
    Levels, score, autopilot, physics — Sonnet ships a whole game loop.Watch at 6:00
  5. 14

    Creativity test: Sol's Wind/Weft sail studio

    The same creative brief sent GPT-6.1 Sol somewhere stranger and more art-directed: Wind/Weft, a sail-design studio where colored thread paths converge into a face before a wind-sovereignty panel with tension readouts and a live mode. The address bar tags model=gpt-6.1-sol. Where Sonnet built a game, Sol built a mood — fewer affordances, stronger identity, the kind of output you show clients to argue models have taste.

    Wind and Weft sail design studio built by GPT-6.1 Sol showing colored thread paths converging into a face outline beside a wind sovereignty panel with tension readouts in live mode.
    Sol's take: an art piece with a settings panel — taste over playability.Watch at 6:25

Which One Should You Pick?

  1. 15

    The routing rule: match the model to the effort setting you actually use

    The decision tree writes itself from the data above. Run Sonnet 5.5 at Low, Medium or High — the settings most coding agents, browser-automation and office workflows use — and GPT-6.1 Sol will likely score higher for about a third of the bill, so pilot it this week. If you deliberately push Sonnet to Max, compare against Opus 5.5, which reaches the same 56 at Xhigh for roughly half the per-task cost. And for the hardest, longest jobs, Opus 5.5 remains the ceiling nobody beat in September.

    Decision tree titled Who should switch routing Claude Sonnet users at low to high effort toward GPT-6.1 Sol for about a third of the bill and max effort users toward Opus with hardest jobs going to Opus 5.5.
    Three arrows, one habit: check your effort setting before you check your loyalty.Watch at 8:58
  2. 16

    Before you migrate: three limits worth knowing

    Three caveats temper the Sol hype. It is not in regular ChatGPT chat yet — you reach it through ChatGPT Work, Codex and the API. Its Ultrafast speed tier is not out either: OpenAI's announcement promises up to 8x faster generation in Codex, while VentureBeat reports Astra's live Ultrafast runs at 6x standard rates, so expect a premium. And these index scores are days old — snapshots move. Offsetting note from launch coverage: OpenAI says Sol's error rate stays within 1.9% of GPT-6 Astra with fewer safety-reviewer circumvention attempts than GPT-6 Sol.

    Limits slide listing not in regular ChatGPT yet, Ultrafast for Sol not out at about six times faster and six times the price, and index scores a day old at most beside TechCrunch and VentureBeat launch quotes.
    Buy the model you can reach today; Ultrafast is a promise, not a product.Watch at 8:30

Frequently asked questions

Keep exploring