GPT-6.1 Sol vs Gemini 3.8: Which Budget AI Should You Actually Use?

OpenAI used DevDay (September 29) to answer Google's cheapest good model with a near-flagship: GPT-6.1 Sol shipped one week after GPT-6 Sol, priced at a fifth of GPT-6 Astra, hours after OpenAI confirmed the pricier GPT-6.1 Astra would not ship. Google's Gemini 3.8 Flash has held the budget crown since September 2 — and it doesn't need a rescue story. This guide puts both launch evidences side by side: sticker prices with their fine print, the benchmark matrix from both labs plus an independent leaderboard, and real builds you can click through. By the end you should be routing tasks between them, not pledging allegiance.

Source & credits

Screenshots in this guide are captured from United Top Tech's public walkthrough video. Every step links back to the exact moment it shows, so you can follow along.

United Top Tech ↗

Two Budget Flagships, Four Weeks Apart

  1. 1

    GPT-6.1 Sol: near-Astra intelligence at a fifth of Astra's price

    OpenAI's DevDay announcement (September 29, 2026) positions GPT-6.1 Sol as an upgrade to the week-old GPT-6 Sol: near GPT-6 Astra's intelligence on agentic coding, computer use and professional work, at one-fifth of Astra's standard token prices. Cached input drops to $0.10 per million tokens — 95% below standard input pricing and half of GPT-6 Sol's cached rate. The swap arrived hours after OpenAI confirmed GPT-6.1 Astra would not ship following internal safety evaluations.

    OpenAI announcement page introducing GPT-6.1 Sol with the headline Near-Astra intelligence for a fifth of the price and a sidebar listing coding, computer use and scientific research.
    The framing OpenAI led with: Astra-class work at 80% off.Watch at 0:30
  2. 2

    Gemini 3.8 Flash: the incumbent, three weeks after 3.7

    Google's announcement post introduces Gemini 3.8 as its best reasoning and coding model yet, built for long-running agentic loops and arriving just three weeks after 3.7 Flash. Two variants shipped on September 2: the standard 3.8 Flash and a separate Gemini 3.8 Flash Cyber for vetted security teams. Unlike OpenAI's replacement scramble, this was a scheduled cadence release — which is exactly the position the Sol launch is attacking.

    Google's post on X introducing Gemini 3.8 with an announcement card that reads Introducing Gemini 3.8 Flash and 3.8 Flash Cyber.
    Google's own pitch: best reasoning and coding model yet.Watch at 0:10

Price Cards: Promo vs Standard

  1. 3

    GPT-6.1 Sol pricing: $2 in, $10 out — same sticker as GPT-6 Sol

    The availability section spells out the deal: GPT-6.1 Sol reached all Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex on day one — not yet in Chat — and developers call it through the OpenAI API as gpt-6.1-sol at $2 per million input tokens, $0.10 cached input and $10 output. An Ultrafast tier with up to 8x faster token generation in Codex was promised for the coming days. Requests above 272K input tokens switch to long-context pricing of $4 and $15.

    Pricing and availability section of OpenAI's GPT-6.1 Sol post listing $2 per million input tokens, $0.10 cached input and $10 output for ChatGPT Work, Codex and the API.
    ChatGPT Work and Codex first, Chat later — read the fine print.Watch at 2:40
  2. 4

    Gemini 3.8 Flash pricing: 75 cents in, $3.75 out — until December 31

    Google AI Studio's model picker shows the promo Gemini 3.8 Flash has run since launch: $0.75 per million input tokens and $3.75 output through December 31, 2026, at all context lengths. From January 1, 2027 the standard rates double to $1.50 and $7.50 — confirmed live on Google's pricing page on September 30, 2026. Even the post-promo price undercuts GPT-6.1 Sol's $2/$10, but Sol's $0.10 cached input undercuts both for repeated agentic calls.

    Google AI Studio model selection panel listing gemini-3.8-flash at $0.75 input and $3.75 output per million tokens through December 31, 2026 with a September 2 release date.
    The promo price is real — and so is the January 1 increase.Watch at 2:20
  3. 5

    Side by side: Flash wins the sticker, Sol wins the cache

    Matthew Berman's comparison table puts the September price landscape in one view: Gemini 3.8 Flash at $0.75/$3.75 (regular $1.50/$7.50) versus Claude Opus 5 at $5/$25, Claude Sonnet 5 at $2/$10, and the GPT-5.6 generation at $2-$4 in and $12-$20 out. GPT-6.1 Sol resets OpenAI's mid-tier to $2/$10 with a $0.10 cached rate. So the price-driven pick between our two headline models is about shape of usage: cold, high-volume tokens favor Flash; cache-heavy agent loops favor Sol.

    Model comparison table scoring Gemini 3.8 Flash at $0.75 input and $3.75 output beside Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5 and the GPT-5.6 tier with DeepSWE and GDPval rows below.
    One table, the whole September price map.Watch at 3:00

The Benchmark Matrix, Both Labs' Numbers Plus One Independent

  1. 6

    OpenAI's chart: Sol 6.1 matches Astra on DeepSWE at a fifth of the cost

    OpenAI's launch chart plots score against cost per task on DeepSWE v1.1: GPT-6.1 Sol's curve climbs into the mid-70s at well under $1 per task, while GPT-6 Astra needs roughly $3-$7.50 for the same plateau. Per the announcement text, 6.1 Sol matches Astra at about one-fifth the cost and beats GPT-6 Sol's best score by 6.4 points at a lower reasoning effort. Independent numbers two steps down test that claim.

    OpenAI DeepSWE cost-per-task chart tracking GPT-6.1 Sol into the mid-70s score range below one dollar while GPT-6 Astra plateaus near 75 percent past three dollars.
    Astra's plateau at Sol pricing — OpenAI's own chart.Watch at 1:25
  2. 7

    Google's table: 3.8 Flash posts 73.7% on the same DeepSWE suite

    Google's eval table scores Gemini 3.8 Flash at 73.7% on DeepSWE v1.1 — effectively tied with GPT-5.6 Sol's 72.7% (Google's chart still benchmarks the previous OpenAI generation — GPT-6.1 Sol only landed on September 29) and a point off Claude Opus 5's 74% — plus 89.4% on Terminal-bench 2.1, 54.9% on HLE-Verified and 59% on OSWorld-2.0 agentic computer use. Terminal-bench 4.0 exposes the trade-off: 19.1% against Opus 5's 51.8%. The footnote repeats the promo deadline: introductory prices expire December 31, 2026, then $1.50/$7.50 applies.

    Google eval table highlighting the Gemini 3.8 Flash column at 73.7 percent on DeepSWE v1.1, 89.4 percent on Terminal-bench 2.1 and 59 percent on OSWorld-2.0 against Claude and GPT-5.6 rivals.
    73.7% on DeepSWE — the same suite where Sol claims Astra parity.Watch at 1:10
  3. 8

    Professional work: GDP.pdf is where Sol 6.1 bullies its price class

    On GDP.pdf, which grades how accurately models answer professional questions from complex PDFs, OpenAI's chart shows GPT-6.1 Sol at 26-32% for under $0.50 per task while Opus 5.5 with fallbacks spends $0.80-$2 for less, and GPT-6 Astra's roughly 30-32% costs about five times more. The announcement claims Sol outscores Opus 5.5 with fallbacks at less than half the cost per task across tested reasoning settings — the yellow curve makes that visible.

    GDP.pdf accuracy-versus-cost scatter showing GPT-6.1 Sol between 26 and 32 percent under fifty cents per task with GPT-6 Sol, GPT-6 Astra and Opus 5.5 with fallbacks plotted to its right.
    Sub-$0.50 per task is the headline of this chart.Watch at 2:00
  4. 9

    AutomationBench: 36.1% at $0.30 per task, hovering near Anthropic

    OpenAI's AutomationBench chart carries a live tooltip on the launch page: GPT-6.1 Sol at max effort costs $0.30 per task and scores 36.1%, with Opus 5.5 (fallbacks) and Fable 5.1 (Opus 5 fallback) plotted above it and GPT-6 Astra ahead of both. The announcement text puts Sol 2.2 points above Opus 5.5 at medium effort for about a third of the cost. If your comparison set is Anthropic's mid-tier rather than Google's, our Claude Sonnet 5.5 vs GPT-6 head-to-head covers the same Sol-versus-Sonnet question in the same format.

    AutomationBench score-versus-cost chart with a tooltip reading GPT-6.1 Sol, Effort Max, cost per task $0.30, score 36.1 percent beside curves for Opus 5.5, Fable 5.1 and GPT-6 Astra.
    The tooltip OpenAI let reviewers hover: 36.1% for $0.30.Watch at 2:28
  5. 10

    The independent view: Flash sits on the efficiency frontier

    Step outside the vendors: Matthew Berman overlays Datacurve's DeepSWE V1.1 leaderboard, where Gemini 3.8 Flash (blue) holds 72-75% quality at the far right of the cost axis — the most efficient cluster, next to Gemini 3.7 Flash and ahead of GPT-5.6 Sol and GPT-5.6 Luna — while Claude Opus 5 and Claude Fable 5 own the expensive left side. Read together with OpenAI's chart, both September models land near 73-75% quality: the real differentiators are price structure, context and surface, not raw score.

    DeepSWE V1.1 scatter from Datacurve plotting Gemini 3.8 Flash at 72 to 75 percent average quality near the zero-cost axis beside Claude Opus 5, Kimi-K3, Grok 4.6 and GPT-5.6 models.
    Vendor-independent DeepSWE: quality converges, cost doesn't.Watch at 7:40

What Gemini 3.8 Actually Builds

  1. 11

    The Cyber variant most buyers can't touch — and Sol has no answer to it

    Gemini 3.8 Flash Cyber tops Google's CyberGym Pass@1 chart at 86.2%, ahead of GPT-5.6 Sol at 83.6%, Mythos 5 at 83.8% and GPT-5.5-Cyber at 85.6% — but it is restricted to vetted defenders through a trusted-defender program, so most teams will never call it. Note what the chart measures: vulnerability discovery in C/C++ codebases. OpenAI's September lineup fields no security-special tier, so in security bake-offs Google simply holds a card Sol cannot match.

    CyberGym Pass@1 bar chart with Gemini 3.8 Flash Cyber leading at 86.2 percent above Gemini 3.5 Flash Cyber, GPT-5.6 Sol, Mythos 5 and GPT-5.5-Cyber for C and C++ vulnerability discovery.
    86.2% on CyberGym — locked behind a vetting program.Watch at 9:00
  2. 12

    Hands-on: a playable Doom-style build inside Antigravity

    Matthew Berman's Antigravity session shows what 3.8 Flash's agentic loop ships: a first-person Doom-style level with a working automap radar, weapon HUD, ammo counter and an enemy to hunt — one prompt in, a game you can actually walk around. This is the workload behind Google's long-running agentic loops pitch, and it competes directly with the Codex sessions OpenAI demos with Sol 6.1.

    Doom-style first-person game built by Gemini 3.8 Flash in Google Antigravity showing a demon in a data-center corridor with an automap radar and an ammo and health HUD.
    One prompt later: enemies, radar, HUD, all live.Watch at 16:10
  3. 13

    Hands-on: an Everest terrain lab that feels like a finished product

    The strongest demo in the test: Gemini 3.8 Flash built a Topographic Lab of Mount Everest in Google Antigravity, complete with a real-time slicing plane, USGS topo layers, a South Col climbing-route profile and an Analyze with Gemini button — the Gemini 3.8 Flash chip sits right in the app header. Data-heavy, tool-using work is where the 1M-token context shows, and it is the category OpenAI points at with GDP.pdf instead.

    Google Antigravity Topographic Lab of Mount Everest built by Gemini 3.8 Flash with a realtime slicing plane panel, USGS topo layer, an 8,848.86 meter summit card and a South Col route profile.
    Data-heavy, tool-using, weirdly polished — classic Flash.Watch at 15:16

Which One Should You Pick?

  1. 14

    Pick by surface: where each model actually lives

    Availability decides half of this matchup. Google ships 3.8 Flash through Antigravity, Google AI Studio and Android Studio for developers, Gemini Enterprise for companies, and AI Pro/Ultra plans in the Gemini app, AI Mode in Search and Google Sheets for consumers. GPT-6.1 Sol answers with ChatGPT Work, Codex and the OpenAI API as gpt-6.1-sol — not yet in Chat — plus an Ultrafast Codex tier. Route by where your work already runs; when in doubt, prototype in both and let per-task cost break the tie.

    Google post on X listing where to find Gemini 3.8 Flash: Antigravity, Google AI Studio and Android Studio for developers, Gemini Enterprise, and AI Pro and Ultra plans in the Gemini app.
    Same model, five doors — check yours before committing.Watch at 2:10

Frequently asked questions

Keep exploring