GPT-6.1 Sol vs Gemini 3.8: Which Budget AI Should You Actually Use?
OpenAI used DevDay (September 29) to answer Google's cheapest good model with a near-flagship: GPT-6.1 Sol shipped one week after GPT-6 Sol, priced at a fifth of GPT-6 Astra, hours after OpenAI confirmed the pricier GPT-6.1 Astra would not ship. Google's Gemini 3.8 Flash has held the budget crown since September 2 — and it doesn't need a rescue story. This guide puts both launch evidences side by side: sticker prices with their fine print, the benchmark matrix from both labs plus an independent leaderboard, and real builds you can click through. By the end you should be routing tasks between them, not pledging allegiance.
Source & credits
Screenshots in this guide are captured from United Top Tech's public walkthrough video. Every step links back to the exact moment it shows, so you can follow along.
United Top Tech ↗Two Budget Flagships, Four Weeks Apart
- 1
GPT-6.1 Sol: near-Astra intelligence at a fifth of Astra's price
OpenAI's DevDay announcement (September 29, 2026) positions GPT-6.1 Sol as an upgrade to the week-old GPT-6 Sol: near GPT-6 Astra's intelligence on agentic coding, computer use and professional work, at one-fifth of Astra's standard token prices. Cached input drops to $0.10 per million tokens — 95% below standard input pricing and half of GPT-6 Sol's cached rate. The swap arrived hours after OpenAI confirmed GPT-6.1 Astra would not ship following internal safety evaluations.

The framing OpenAI led with: Astra-class work at 80% off.Watch at 0:30 - 2
Gemini 3.8 Flash: the incumbent, three weeks after 3.7
Google's announcement post introduces Gemini 3.8 as its best reasoning and coding model yet, built for long-running agentic loops and arriving just three weeks after 3.7 Flash. Two variants shipped on September 2: the standard 3.8 Flash and a separate Gemini 3.8 Flash Cyber for vetted security teams. Unlike OpenAI's replacement scramble, this was a scheduled cadence release — which is exactly the position the Sol launch is attacking.

Google's own pitch: best reasoning and coding model yet.Watch at 0:10
Price Cards: Promo vs Standard
- 3
GPT-6.1 Sol pricing: $2 in, $10 out — same sticker as GPT-6 Sol
The availability section spells out the deal: GPT-6.1 Sol reached all Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex on day one — not yet in Chat — and developers call it through the OpenAI API as gpt-6.1-sol at $2 per million input tokens, $0.10 cached input and $10 output. An Ultrafast tier with up to 8x faster token generation in Codex was promised for the coming days. Requests above 272K input tokens switch to long-context pricing of $4 and $15.

ChatGPT Work and Codex first, Chat later — read the fine print.Watch at 2:40 - 4
Gemini 3.8 Flash pricing: 75 cents in, $3.75 out — until December 31
Google AI Studio's model picker shows the promo Gemini 3.8 Flash has run since launch: $0.75 per million input tokens and $3.75 output through December 31, 2026, at all context lengths. From January 1, 2027 the standard rates double to $1.50 and $7.50 — confirmed live on Google's pricing page on September 30, 2026. Even the post-promo price undercuts GPT-6.1 Sol's $2/$10, but Sol's $0.10 cached input undercuts both for repeated agentic calls.

The promo price is real — and so is the January 1 increase.Watch at 2:20 - 5
Side by side: Flash wins the sticker, Sol wins the cache
Matthew Berman's comparison table puts the September price landscape in one view: Gemini 3.8 Flash at $0.75/$3.75 (regular $1.50/$7.50) versus Claude Opus 5 at $5/$25, Claude Sonnet 5 at $2/$10, and the GPT-5.6 generation at $2-$4 in and $12-$20 out. GPT-6.1 Sol resets OpenAI's mid-tier to $2/$10 with a $0.10 cached rate. So the price-driven pick between our two headline models is about shape of usage: cold, high-volume tokens favor Flash; cache-heavy agent loops favor Sol.

One table, the whole September price map.Watch at 3:00
The Benchmark Matrix, Both Labs' Numbers Plus One Independent
- 6
OpenAI's chart: Sol 6.1 matches Astra on DeepSWE at a fifth of the cost
OpenAI's launch chart plots score against cost per task on DeepSWE v1.1: GPT-6.1 Sol's curve climbs into the mid-70s at well under $1 per task, while GPT-6 Astra needs roughly $3-$7.50 for the same plateau. Per the announcement text, 6.1 Sol matches Astra at about one-fifth the cost and beats GPT-6 Sol's best score by 6.4 points at a lower reasoning effort. Independent numbers two steps down test that claim.

Astra's plateau at Sol pricing — OpenAI's own chart.Watch at 1:25 - 7
Google's table: 3.8 Flash posts 73.7% on the same DeepSWE suite
Google's eval table scores Gemini 3.8 Flash at 73.7% on DeepSWE v1.1 — effectively tied with GPT-5.6 Sol's 72.7% (Google's chart still benchmarks the previous OpenAI generation — GPT-6.1 Sol only landed on September 29) and a point off Claude Opus 5's 74% — plus 89.4% on Terminal-bench 2.1, 54.9% on HLE-Verified and 59% on OSWorld-2.0 agentic computer use. Terminal-bench 4.0 exposes the trade-off: 19.1% against Opus 5's 51.8%. The footnote repeats the promo deadline: introductory prices expire December 31, 2026, then $1.50/$7.50 applies.

73.7% on DeepSWE — the same suite where Sol claims Astra parity.Watch at 1:10 - 8
Professional work: GDP.pdf is where Sol 6.1 bullies its price class
On GDP.pdf, which grades how accurately models answer professional questions from complex PDFs, OpenAI's chart shows GPT-6.1 Sol at 26-32% for under $0.50 per task while Opus 5.5 with fallbacks spends $0.80-$2 for less, and GPT-6 Astra's roughly 30-32% costs about five times more. The announcement claims Sol outscores Opus 5.5 with fallbacks at less than half the cost per task across tested reasoning settings — the yellow curve makes that visible.

Sub-$0.50 per task is the headline of this chart.Watch at 2:00 - 9
AutomationBench: 36.1% at $0.30 per task, hovering near Anthropic
OpenAI's AutomationBench chart carries a live tooltip on the launch page: GPT-6.1 Sol at max effort costs $0.30 per task and scores 36.1%, with Opus 5.5 (fallbacks) and Fable 5.1 (Opus 5 fallback) plotted above it and GPT-6 Astra ahead of both. The announcement text puts Sol 2.2 points above Opus 5.5 at medium effort for about a third of the cost. If your comparison set is Anthropic's mid-tier rather than Google's, our Claude Sonnet 5.5 vs GPT-6 head-to-head covers the same Sol-versus-Sonnet question in the same format.

The tooltip OpenAI let reviewers hover: 36.1% for $0.30.Watch at 2:28 - 10
The independent view: Flash sits on the efficiency frontier
Step outside the vendors: Matthew Berman overlays Datacurve's DeepSWE V1.1 leaderboard, where Gemini 3.8 Flash (blue) holds 72-75% quality at the far right of the cost axis — the most efficient cluster, next to Gemini 3.7 Flash and ahead of GPT-5.6 Sol and GPT-5.6 Luna — while Claude Opus 5 and Claude Fable 5 own the expensive left side. Read together with OpenAI's chart, both September models land near 73-75% quality: the real differentiators are price structure, context and surface, not raw score.

Vendor-independent DeepSWE: quality converges, cost doesn't.Watch at 7:40
What Gemini 3.8 Actually Builds
- 11
The Cyber variant most buyers can't touch — and Sol has no answer to it
Gemini 3.8 Flash Cyber tops Google's CyberGym Pass@1 chart at 86.2%, ahead of GPT-5.6 Sol at 83.6%, Mythos 5 at 83.8% and GPT-5.5-Cyber at 85.6% — but it is restricted to vetted defenders through a trusted-defender program, so most teams will never call it. Note what the chart measures: vulnerability discovery in C/C++ codebases. OpenAI's September lineup fields no security-special tier, so in security bake-offs Google simply holds a card Sol cannot match.

86.2% on CyberGym — locked behind a vetting program.Watch at 9:00 - 12
Hands-on: a playable Doom-style build inside Antigravity
Matthew Berman's Antigravity session shows what 3.8 Flash's agentic loop ships: a first-person Doom-style level with a working automap radar, weapon HUD, ammo counter and an enemy to hunt — one prompt in, a game you can actually walk around. This is the workload behind Google's long-running agentic loops pitch, and it competes directly with the Codex sessions OpenAI demos with Sol 6.1.

One prompt later: enemies, radar, HUD, all live.Watch at 16:10 - 13
Hands-on: an Everest terrain lab that feels like a finished product
The strongest demo in the test: Gemini 3.8 Flash built a Topographic Lab of Mount Everest in Google Antigravity, complete with a real-time slicing plane, USGS topo layers, a South Col climbing-route profile and an Analyze with Gemini button — the Gemini 3.8 Flash chip sits right in the app header. Data-heavy, tool-using work is where the 1M-token context shows, and it is the category OpenAI points at with GDP.pdf instead.

Data-heavy, tool-using, weirdly polished — classic Flash.Watch at 15:16
Which One Should You Pick?
- 14
Pick by surface: where each model actually lives
Availability decides half of this matchup. Google ships 3.8 Flash through Antigravity, Google AI Studio and Android Studio for developers, Gemini Enterprise for companies, and AI Pro/Ultra plans in the Gemini app, AI Mode in Search and Google Sheets for consumers. GPT-6.1 Sol answers with ChatGPT Work, Codex and the OpenAI API as gpt-6.1-sol — not yet in Chat — plus an Ultrafast Codex tier. Route by where your work already runs; when in doubt, prototype in both and let per-task cost break the tie.

Same model, five doors — check yours before committing.Watch at 2:10
Frequently asked questions
Keep exploring
- ChatGPT tool profile: plans, models and setup
- Gemini tool profile: plans, models and setup
- GPT-6 vs Gemini 3: the flagship head-to-head
- Claude Sonnet 5.5 vs Gemini 3.8 Flash: the other budget matchup
- How to use GPT-6.1 Sol: the full getting-started guide
- Claude Sonnet 5.5 vs GPT-6: the Sonnet side of this story

