Claude Sonnet 5.5 vs Gemini 3.8 Flash: Which One Should You Actually Use?

September 2026 gave budget-conscious builders two very different mid-tier models. Google's Gemini 3.8 Flash (September 2) is the third Flash release in six weeks — a fast, dirt-cheap workhorse that Google says now codes and runs agents like a much pricier model. Anthropic's Claude Sonnet 5.5 (September 28) keeps Sonnet pricing but claims near-Opus 5.5 judgment for half of Opus's price, with a reported Artificial Analysis Intelligence Index of 56 — three points above GPT-6 Astra. This walkthrough compares them step by step, with real price tables, benchmark charts and hands-on sessions from five independent testers: what each costs, where each wins, and which one fits your actual work.

Source & credits

Screenshots in this guide are captured from Bijan Bowen's public walkthrough video. Every step links back to the exact moment it shows, so you can follow along.

Bijan Bowen ↗

The 60-Second Verdict

  1. 1

    Two September launches, 26 days apart

    Google shipped Gemini 3.8 Flash on September 2 — its third Flash release in roughly six weeks, and the announcement is explicit about the formula: the same speed and low cost as 3.7, with the biggest gains in coding and agentic work. Anthropic answered on September 28 with Claude Sonnet 5.5, a mid-tier model positioned as a faster, cheaper companion to Opus 5.5. So the matchup is really budget speed versus near-flagship judgment — and both camps want to be your default daily model.

    Google's announcement post introduces Gemini 3.8 Flash as its third Flash release in six weeks, promising the same speed and low cost as 3.7 Flash with the biggest gains in coding and agents.
    Three Flash releases in six weeks set up Google's budget pitch.Watch at 0:38
  2. 2

    The short answer: quality per dollar

    On an independent long-horizon software-engineering leaderboard recorded in The Average Tech Dad's test, Gemini 3.8 Flash and GPT-6 Astra both score 74% — but Gemini's average cost is $2.36 per task against Astra's $6.52, with Claude Opus 5 at $11.84. That is the Flash value story in one screen. Sonnet 5.5 answers differently: Anthropic says it needs far fewer tokens per task than Sonnet 5, cutting cost per task by up to 30% while running about 30% faster, and Artificial Analysis reportedly scores it 56 on its Intelligence Index — three points above GPT-6 Astra. Gemini wins the price chart; Sonnet 5.5 is the cheapest route to near-flagship judgment.

    Independent agent leaderboard shows Gemini 3.8 Flash scoring 74 percent for $2.36 per task while GPT-6 Astra spends $6.52 and Claude Opus 5 spends $11.84 for the same score.
    Same 74% score: Flash pays $2.36, Astra $6.52, Opus 5 $11.84.Watch at 1:00

Price Cards and Fine Print

  1. 3

    Sonnet 5.5 pricing: unchanged sticker, cheaper tasks

    Anthropic kept Sonnet pricing exactly where it was: $2 per million input tokens and $10 per million output, with cache reads at $0.20 and cache writes at $2.50 — exactly half of Opus 5.5 across the board, as the cost-and-speed table in Fahd Mirza's walkthrough shows. The real saving is behavioral: Anthropic says the model typically needs far fewer tokens to finish the same task, which is why it claims up to 30% lower cost per task than Sonnet 5 plus roughly 30% faster output. One warning from the same video: at $10 per million output tokens, an unattended agent can still run up a surprising bill — the demo's trampoline game burned roughly $25-30.

    Anthropic's cost and speed table lists Claude Sonnet 5.5 at $2 input and $10 output per million tokens with cache reads of $0.20, half of Claude Opus 5.5's rates.
    Sonnet 5.5 keeps Sonnet 5's price card: $2 in, $10 out.Watch at 1:04
  2. 4

    Gemini 3.8 Flash pricing: 75 cents in, $3.75 out — for now

    Google's launch table prices Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output — an introductory 50% discount that expires on December 31, 2026, after which both figures double to $1.50 and $7.50, per the footnote on Google's own eval table. Even the regular rate undercuts every Claude column on the same table: Claude Sonnet 5 sits at $2 and $10, Claude Opus 5 at $5 and $25. The DeepSWE v1.1 row makes the pitch concrete: 73.7% for 3.8 Flash versus 74.0% for Claude Opus 5 — near-identical long-horizon engineering output at a fraction of the listed price.

    Google's launch table compares Gemini 3.8 Flash pricing of $0.75 input and $3.75 output per million tokens against Claude Opus 5, Claude Sonnet 5 and GPT-5.6 tiers.
    Flash undercuts every Claude and GPT column on Google's own table.Watch at 1:10
  3. 5

    What you get at each price: context and capabilities

    The Gemini API card for gemini-3.8-flash reads like a workhorse spec sheet: 1,048,576 input tokens, 65,536 output tokens, and text as the only output type — it accepts text, image, video, audio and PDF but will never generate a picture or a clip back. Search grounding, code execution, function calling and URL context are all supported, computer use is in preview, and thinking ships in low, medium and high levels with medium as the default. Sonnet 5.5 counters with roughly 262K context, genuine computer-use strength and document polish that Anthropic demonstrates with a 10-slide earnings review two reviewers called submission-ready.

    Gemini API documentation for gemini-3.8-flash lists a 1,048,576 input token limit, 65,536 output tokens, text-only output and low, medium and high thinking levels.
    The exact numbers: 1,048,576 tokens in, 65,536 out, text only.Watch at 8:30

Where Claude Sonnet 5.5 Pulls Ahead

  1. 6

    It really does sit on Opus 5.5's shoulders

    Anthropic's own release table is the strongest evidence for the near-Opus-at-half-price claim: on FrontierCode 1.1 Sonnet 5.5 reaches 52.1% at extra-high effort against Opus 5.5's 54.4%, CursorBench 4.0 puts them at 55.5% versus 57.8%, and knowledge-work scores land within a point or two — 1844 versus 1846 on GDPval-AA. Computer use is nearly a tie at 80.1% versus 81.8% on OSWorld 2.1, at half the token price. These are vendor figures, and independent scorecards were still landing as this guide was written, but the shape is consistent: for everyday engineering, Sonnet 5.5 behaves like Opus 5.5 with a smaller bill.

    Anthropic's release table puts Claude Sonnet 5.5 within a few points of Claude Opus 5.5 on FrontierCode, CursorBench, OSWorld and knowledge-work benchmarks at half the price.
    Anthropic's own table: Sonnet 5.5 shadows Opus 5.5 row by row.Watch at 2:50
  2. 7

    The effort dial is where the value hides

    Fahd Mirza's Terminal-Bench 4.0 accuracy-versus-cost chart explains how to actually save money with Sonnet 5.5. At low and medium effort the blue line already beats everything Sonnet 5 could do — Anthropic claims its best low-or-medium effort results beat Sonnet 5's top scores at about a tenth of the cost per task. Push to extra-high or max and the model keeps climbing, overtaking Opus 5.5 entirely around the max setting while Opus plateaus and dips. His advice matches the chart: live on medium, save high for work that justifies it, because high effort drains both money and rate limits quickly.

    Terminal-Bench 4.0 accuracy-versus-cost chart traces Claude Sonnet 5.5 climbing past Claude Opus 5.5 as effort rises from low to maximum.
    Max effort is where Sonnet 5.5 crosses Opus 5.5.Watch at 3:45
  3. 8

    Docs, slides and polish: the quiet Sonnet specialty

    Anthropic positions Sonnet 5.5 as the everyday line for well-defined work: bug fixing, spreadsheets, presentations and documents. Its launch example — a 10-slide earnings review that two reviewers judged ready to send unchanged — is what separates it from budget models. Early customers quoted at launch back the polish claim: Box reports 2.4x faster output at higher accuracy, Zendesk processed tickets 20% faster, and Slack saw about 14% fewer output tokens on comparable writing. It is also the first Sonnet to carry Opus-grade cyber safeguards, so routine development is unaffected while high-risk security requests get flagged.

    Sonnet 5.5 explainer slide maps the model's best-use stops: bug fixing, docs and slides, UI polish, computer use and support at scale.
    Five stops where Anthropic says Sonnet 5.5 shines.Watch at 2:42

Where Gemini 3.8 Flash Pulls Ahead

  1. 9

    Near-frontier scores at a fraction of the bill

    Google's eval table is aggressive in a way budget models usually are not: Gemini 3.8 Flash posts 86.2% on CharXiv Reasoning, 87.8% agentic on LV-Bench, 54.9% on HLE-Verified and 88.8% on BioMysteryBench — trading leads with Claude Opus 5 row by row while costing a fraction as much per output token. In Google's own tests the release jumped Terminal-Bench from 82% to 91% and DeepSWE from 65% to 74%, gains the company attributes to greater diligence: extra reasoning steps and iterative tool calls, especially at higher effort. Treat vendor numbers with caution, but Bijan Bowen's hands-on reached the same verdict — clearly better than 3.7, very fast and fairly cheap.

    Google's eval table scores Gemini 3.8 Flash at 86.2 percent on CharXiv Reasoning, 59 percent on OSWorld-2.0 and 88.8 percent on BioMysteryBench beside Claude and GPT-5.6 columns.
    Google's eval table: near-frontier scores at budget pricing.Watch at 1:35
  2. 10

    One prompt in, a playable 3D world out

    The Average Tech Dad handed Gemini 3.8 Flash a long-running brief in Google's Antigravity IDE — on High effort with a Pro subscription — and got back The Constituency, a complete 3D strategy game: procedurally seeded maps, sixteen turns, a budget in local currency, satisfaction and scrutiny meters, a build palette from boreholes to water works, and an isometric camera. He notes free-tier users should still get some mileage, just not a full game per session. That is the Flash trade in miniature: less refined judgment than a flagship, but fast enough to iterate and cheap enough to experiment.

    The Constituency, a 3D strategy game built by Gemini 3.8 Flash in Antigravity, shows its budget bar, build palette, scrutiny meter and isometric map.
    A 3D governance strategy game, budget AI edition.Watch at 5:35
  3. 11

    Whole apps from one brief — synthwave included

    Bijan Bowen's browser-OS test shows the same pattern from another angle: one prompt to Antigravity produced an entire synthwave desktop in the browser, complete with a windowed mail app, a code editor, procedural wallpapers and a built-in Cyber-Doom 3D neon arena shooter with wave counters and a score readout. He jokes that every Flash generation loves synthwave — the aesthetic has survived several model generations. The takeaway for buyers: Gemini 3.8 Flash is fast and cheap enough that one-brief creative projects stop being a luxury, though he still found rough edges, like a watch-store site whose gears spun on the wrong axis.

    A synthwave browser OS built by Gemini 3.8 Flash runs its built-in Cyber-Doom 3D neon arena game with a wave counter and score readout in the browser.
    The browser OS Gemini built — neon arena included.Watch at 5:30

Head-to-Head: Coding, Documents, Cost, Agents

  1. 12

    Coding: polish and judgment versus speed and price

    Real-world coding is where the philosophical split shows. Sonnet 5.5 works like a senior pair-programmer: in Fahd Mirza's Claude Code session it planned a Blender-plus-Godot trampoline game, caught inconsistencies in the brief's own numbers, and shipped a playable build — for roughly $25-30 of tokens. Gemini 3.8 Flash codes like a sprinter: Bijan Bowen's C++ skate game improved dramatically when the model was allowed to self-review, and Gemini finished a robot-arm pick-and-place task in 36 minutes — faster than GPT-6 Astra's 40 on the same rig. If you measure accepted-output quality per hour of review, pick Sonnet 5.5; if you measure iterations per dollar, pick Gemini.

    Claude Code terminal runs a Claude Sonnet 5.5 session building a 3D trampoline game, showing token usage and 27 tokens per second in the status bar.
    A real Sonnet 5.5 coding run in Claude Code.Watch at 0:40
  2. 13

    Long documents: similar windows, different instincts

    Both models swallow very long inputs — Gemini 3.8 Flash formally lists 1,048,576 input tokens, and BitBiasedAI's test pasted an entire two-hour meeting transcript and asked for every task, owner and deadline, which the model handled without losing the thread between minute 10 and minute 90. Its multimodal reading is a genuine advantage: feed it an ad-dashboard screenshot and it returns a metric-by-metric diagnosis like the one in this frame, naming CTR, conversions and cost per conversion with likely causes. Sonnet 5.5 reads long context too (about 262K on standard settings) but earns its keep on the output side — polishing decks and documents rather than just summarizing them.

    Gemini 3.8 Flash reads an uploaded ad dashboard screenshot and lists CTR, conversions and cost per conversion as the underperforming metrics with likely causes.
    Paste a dashboard screenshot, get a metric-by-metric diagnosis.Watch at 9:05
  3. 14

    Speed and cost: Flash wins the sticker, Sonnet wins per unit of judgment

    The same independent leaderboard that opened this guide makes the cost gap concrete: 74% quality cost $2.36 with Gemini 3.8 Flash, $6.46 with GPT-5.6 Sol, $6.52 with GPT-6 Astra and $11.84 with Claude Opus 5. Gemini is also blunt-force fast — reviewers repeatedly describe Flash output as blazingly fast. Sonnet 5.5 never wins that chart on price, but Anthropic's efficiency claims — up to 30% cheaper per task than Sonnet 5, about 30% faster, roughly a tenth of the cost at low-medium effort for better-than-Sonnet-5 scores — reframe the comparison: you are paying for fewer review cycles, not just fewer dollars.

    Agent benchmark scoreboard lists Gemini 3.8 Flash at $2.36 average cost beside GPT-5.6 Sol at $6.46, GPT-6 Astra at $6.52 and Claude Opus 5 at $11.84 for similar 74 percent scores.
    What equal quality costs across frontier models.Watch at 1:30
  4. 15

    Agents: Gemini's vision tricks versus Sonnet's guardrails

    For autonomous work the two play different sports. Gemini 3.8 Flash drove a physical robot arm through a pick-and-place task in 36 minutes using a deliberately bad camera angle, writing its own vision pipeline and mission script — and Google ships agent tooling as a first-class use case, with computer use in preview and search grounding built in. Sonnet 5.5 answers with reliability plumbing: Opus-grade cyber safeguards, near-Opus computer-use scores (80.1% on OSWorld 2.1), day-one availability through AWS and, reportedly, GitHub Copilot, plus the token discipline that keeps long agent loops from overspending. Choose by failure mode: Gemini when perception matters, Sonnet when the loop runs unattended.

    A blue robotic arm lifts an orange toy car while Gemini 3.8 Flash's Antigravity session writes the vision pipeline and pick-and-place script on the monitor behind.
    36 minutes, one robotic arm, zero human hands.Watch at 16:20

Which One Should You Pick?

  1. 16

    Pick by role, not by brand

    The decision gets easy once you translate it into jobs. Pick Claude Sonnet 5.5 for everyday engineering, bug fixing, client deliverables, documents and slides, computer-use workflows and any unattended loop where a wrong answer is expensive — it is the closest a mid-tier model gets to Opus 5.5 judgment at half price. Pick Gemini 3.8 Flash for high-volume tasks, quick drafts, search-grounded answers, multimodal reading of screenshots and PDFs, and creative one-prompt builds where iterations cost cents. And ignore the Gemini 4 chatter for now: Gemini 4 is not released — only leaks — so 3.8 Flash is Google's current mid-tier and the fair comparison on this page. Many teams will run both and route by task.

    Model chooser slide splits tasks between Sonnet 5.5 for clear well-scoped work and Opus 5.5 for complex open-ended jobs where mistakes are expensive.
    Two lines, two jobs: route tasks, don't marry brands.Watch at 4:02

Frequently asked questions

Keep exploring