Skip to content

GPT-6 vs Claude Opus 5.5 vs Gemini 3.8: Which AI Model Wins? (2026)

Matt Li By Matt Li Co-Founder and Director 12 min read
TL;DR: Claude Opus 5.5 leads the independent Artificial Analysis Intelligence Index at 57.6, ahead of GPT-6 Astra at 52.7 and Gemini 3.8 Flash at 40.9, and its API costs $4 in and $20 out per million tokens against Astra's $10 and $50. GPT-6 Astra is the pick for computer use and hard maths, where OpenAI's computer-use results and Epoch AI's independent maths runs put it first. Gemini 3.8 Flash is the fast, cheap option at $0.75 and $3.75, because Google has not released a Gemini 3.8 Pro.

On ARC-AGI-3, GPT-6 Astra solved 62.7% of the puzzles through the ARC Prize Foundation's standard harness and 99.9% through the adapter OpenAI supplied. OpenAI's launch post quotes the second number.

Anthropic released Claude Opus 5.5 on September 22, 90 minutes before OpenAI's cheaper GPT-6 Sol and Luna, and Google's newest model this month is still a Flash. I read all three launch posts, the Opus 5.5 system card and the price sheets; the scores below come from published evals, not my own runs.

What decides the pick
  1. 1Anthropic reports 66.4% for Opus 5.5 on Terminal-Bench 4.0. Artificial Analysis measured 59.6%, level with Astra's 59.1%.
  2. 2In LMArena's blind text vote, Opus 5.5 ranks 1st, Gemini 3.8 Flash 10th and GPT-6 Astra 26th.
  3. 3Every GPT-6 prompt over 272K tokens is billed at twice the input rate and 1.5 times the output rate.
  4. 4Gemini 3.8 Flash's price doubles to $1.50 and $7.50 on January 1, 2027.

Which Model Is Best Right Now?

Claude Opus 5.5, on the two independent boards that have scored all three. The Artificial Analysis leaderboard put it first on September 29 with 57.6, and it leads four of the five evals below. Its cheaper sibling, Claude Sonnet 5.5, sits second at 56.0 and has the best Terminal-Bench 4.0 run of any model tested.

Heatmap of Artificial Analysis scores. Intelligence Index v4.3: Claude Opus 5.5 57.6, Claude Sonnet 5.5 56.0, GPT-6 Astra 52.7, GPT-6 Sol 47.5, Gemini 3.8 Flash 40.9, GPT-6 Luna 37.3. Opus 5.5 leads HLE, SciCode, long-context retrieval and MMMU; Sonnet 5.5 leads Terminal-Bench 4.0 at 63.6.

GPT-6 Astra scores 52.7 and Gemini 3.8 Flash 40.9. Astra's gap is widest on SciCode and Humanity's Last Exam; on Terminal-Bench 4.0 it trails Opus 5.5 by half a point. LMArena's text leaderboard agrees on the order: Opus 5.5 rates 1,509, Gemini 3.8 Flash 1,492 and Astra 1,478, though Opus has only 2,307 votes so far.

Which index version. OpenAI's Astra post quotes version 4.1.1 of the Artificial Analysis index, where Astra scores 61.2 and Claude Fable 5.1 scores 65.7. The live board is version 4.3, a harder scale, so scores from the two versions do not line up.

Release Dates, Prices and Context Windows

All three flagships read about a million tokens, but only Gemini 3.8 Flash takes audio and video in. Astra and Opus 5.5 accept text and images and return text.

ModelReleasedInput / output / cache read, per 1MContext / max outputConsumer plans
GPT-6 AstraSep 3, 2026$10 / $50 / $11,050,000 / 128KChatGPT Plus, Pro, Business, Enterprise
Claude Opus 5.5Sep 22, 2026$4 / $20 / $0.201M / 128KClaude Pro, Max, Team, Enterprise
Gemini 3.8 FlashSep 2, 2026$0.75 / $3.75 / $0.0751,048,576 / 65,536Google AI Pro and Ultra
GPT-6 SolSep 22, 2026$2 / $10 / $0.201,050,000 / 128KChatGPT Work and Codex, most paid plans
GPT-6 LunaSep 22, 2026$0.10 / $0.50 / $0.011,050,000 / 128KAlso Free and Go
Claude Sonnet 5.5Sep 28, 2026$2 / $10 / $0.201M / 128KAll Claude plans, including Free

Prices come from the OpenAI API pricing page, Claude's pricing page and Google's Gemini API pricing, read on September 29. The six releases landed inside four weeks:

Sep 2, 2026
Gemini 3.8 Flash and 3.8 Flash Cyber, Google's third Flash release in six weeks
Sep 3, 2026
GPT-6 Astra to a limited set of organisations, then ChatGPT and the API over the following days
Sep 22, 2026
Claude Opus 5.5, then GPT-6 Sol and Luna about 90 minutes later at half the GPT-5.6 price
Sep 28, 2026
Claude Sonnet 5.5 at Sonnet 5's price; Haiku 5.5 is promised "in the coming weeks"
Jan 1, 2027
Gemini 3.8 Flash moves from its introductory price to $1.50 and $7.50

GPT-6 Astra: Computer Use and Maths

OpenAI's GPT-6 Astra launch page, with the September 22 update adding GPT-6 Sol and Luna and the rollout to ChatGPT Plus, Pro, Business and Enterprise.
167 Epoch Capabilities Index, 1st of 25192.7% ScreenSpot-Pro, no tools291 s median wait at max effort

GPT-6 Astra is for agents that drive a desktop or browser, and for maths and science work. On OpenAI's launch table it scores 72.6% on the offline OSWorld 2.0 set in about 40 minutes per task, against 65.7% in 75 minutes for GPT-5.6 Sol.

The maths lead has outside backing. Epoch AI scores it 98% on FrontierMath Tier 4 and calls that tier saturated, and on its new FrontierMath Erdős set of 68 unsolved problems Astra scores 3%, where Claude Fable 5.1 scores 0%. Epoch had not published an index score for Opus 5.5 by September 29.

ARC-AGI-3 needs a footnote. The ARC Prize Foundation verified both results below, and the gap comes from what the harness lets the model carry between turns. OpenAI's adapter keeps its opaque reasoning state and compacts long conversations; the standard harness allows visible notes only.

Arrow chart of GPT-6 Astra on ARC-AGI-3 by harness: max effort 62.7% standard vs 98.6% provider adapter, xhigh 59.3 vs 98.4, high 54.8 vs 99.9, medium 38.6 vs 98.4, low 17.5 vs 98.0.

"GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness."

ARC Prize Foundation, September 3, 2026
  • Best for: computer-use agents, research maths and science, long-document retrieval inside OpenAI's own tests (96.3% on 8-needle MRCR at 512K to 1M tokens).
  • Where it struggles: price and chat preference. It is 26th on LMArena text, and at max effort Artificial Analysis waits a median 291 seconds for the first answer token.
  • Price: $10 in, $50 out, $1 cached; Fast mode doubles both the speed and the price.

Claude Opus 5.5: Coding and Knowledge Work at a Lower Price

Anthropic's Claude Opus 5.5 launch page, dated September 22, 2026, with sections for introduction, performance and cost, safety and availability.
89.9% SWE-bench Pro, Anthropic's run1,846 GDPval-AA v2.1 Elo1st in LMArena's code arena

Claude Opus 5.5 is for coding agents, long migrations and document-heavy office work. Anthropic's launch post says it costs 40% less to run than Opus 5 on typical workloads: tokens are 20% cheaper, cache reads 60% cheaper, and it spends fewer tokens per task.

The Opus 5.5 system card reports 89.9% on SWE-bench Pro, 93.9% on SWE-bench Multilingual and 81.8% partial credit on OSWorld 2.0. OpenAI published no SWE-bench Pro score for Astra. On five evals where both labs do publish a number, Anthropic's own table shows a close race:

Radar chart of vendor-reported scores. Claude Opus 5.5 vs GPT-6 Astra: HLE with tools 67.7 vs 57.2, FrontierCode Main 54.4 vs 53.3, Terminal-Bench Science 58.7 vs 64.6, AutomationBench 40.0 vs 41.4, HealthBench Professional 65.6 vs 63.4.

Astra wins Terminal-Bench Science and AutomationBench; Opus 5.5 wins Humanity's Last Exam with tools by 10.5 points. The independent LMArena code arena rates Opus 5.5 at 1,827 and Astra at 1,792, with Gemini 3.8 Flash 28th at 1,580.

Anthropic ran its benchmarks with production safeguards on. If a cyber task trips them, Opus 4.8 answers instead, so those scores likely understate the raw model. Our Claude Fable vs Opus vs Sonnet guide covers where Fable 5.1 still fits above it.

  • Best for: agentic coding, codebase-wide refactors, finance and legal drafting, chat quality.
  • Where it struggles: latency at max effort, where Artificial Analysis clocks a median 681 seconds before the first answer token. It has no thinking-off mode.
  • Price: $4 in, $20 out, $0.20 cache reads, $5 cache writes; Fast mode is $8 and $40.

Gemini 3.8 Flash: Google's Only 3.8 Model

Google's blog post Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, dated September 2, 2026.
239 output tokens a second73.7% DeepSWE v1.1, Google's runMarch 2026 knowledge cutoff

Gemini 3.8 Flash is for high-volume agents and anything with video or audio input. Google's newest Pro model is still Gemini 3.1 Pro Preview from February, at $2 and $12, so the "Gemini 3.8" in any comparison this month is a Flash-tier model priced under a fifth of Opus 5.5's input rate.

The Gemini 3.8 Flash model card puts it at 73.7% on DeepSWE v1.1, against 74.0% for Claude Opus 5, and 59.0% on OSWorld 2.0. Google built it on Gemini 3.7 Flash rather than a new base. Its weak spot is Terminal-Bench 4.0 at 19.1%, a figure the public leaderboard and Artificial Analysis both confirm.

"These performance gains stem from a core design choice: 3.8 Flash works harder."

Tulsee Doshi and Raluca Ada Popa, Google, September 2, 2026

Working harder means more thinking tokens, and Google tells efficiency-first users to stay on Gemini 3.7 Flash, which costs the same. Our guide to every Gemini model tracks the Flash and Pro lines, and our Gemini statistics page covers its user numbers.

  • Best for: fast, cheap agent loops, video and audio understanding (87.8% on LVBench in Google's agentic run), Google Workspace users.
  • Where it struggles: hard terminal tasks, and a price that doubles in January.
  • Price: $0.75 in, $3.75 out (thinking tokens included), $0.075 cached.

Where Vendor Scores and Independent Boards Disagree

Anthropic printed 66.4% for Opus 5.5 and 70.6% for Sonnet 5.5 on Terminal-Bench 4.0; Artificial Analysis's own runs gave 59.6% and 63.6%. OpenAI's and Google's claims survive the rerun within a point. Anthropic's own footnote gives Opus 5.5 a standard error of plus or minus 2.6 points, and it scored Opus at xhigh effort.

Dumbbell chart of Terminal-Bench 4.0, vendor-reported vs Artificial Analysis: Claude Sonnet 5.5 70.6 vs 63.6, Claude Opus 5.5 66.4 vs 59.6, GPT-6 Astra 57.9 vs 59.1, Gemini 3.8 Flash 19.1 vs 19.7.

The official Terminal-Bench leaderboard lists Astra at 58.2% with Codex, but had no Opus 5.5 or Sonnet 5.5 entry on September 29. Harness matters too: the board grades Claude models in Claude Code, Astra in Codex and Gemini in mini-SWE-agent.

The labs also print different scores for the same rival. On OSWorld 2.0, Google's model card lists Claude Opus 5 at 75.4%, while OpenAI's table lists it at 70.2% on the offline subset it reran. SWE-bench Verified cannot settle the coding question: the SWE-bench leaderboard has no entries newer than February 2026. METR tested Opus 5.5 before launch, but it has not published a time horizon for either flagship.

Speed and Time to First Token

Gemini 3.8 Flash answers in about 22 seconds at high effort. Opus 5.5 at max effort thinks for a median 11 minutes before it writes, then streams faster than Astra.

Split bar chart of median seconds to first answer token and output tokens per second: Gemini 3.8 Flash 22 s and 239, GPT-6 Luna 94 s and 154, Claude Sonnet 5.5 328 s and 139, Claude Opus 5.5 681 s and 94, GPT-6 Sol 164 s and 85, GPT-6 Astra 291 s and 63.

Those waits shrink at lower settings. On the same board, Opus 5.5 at xhigh effort takes a median 136 seconds and Astra at xhigh 124 seconds. Both vendors sell a Fast mode: OpenAI's gives Astra up to twice the speed at twice the price, and Anthropic's gives Opus 5.5 up to 2.5 times the speed at twice the price.

API Price per Million Tokens

Astra's output tokens cost 2.5 times Opus 5.5's, but Astra spends fewer of them. Artificial Analysis ran its full index at each model's top setting, and the bill per task came out lower for Astra:

$5.98
Claude Opus 5.5, max effort
$3.26
GPT-6 Astra, max effort
$1.24
Gemini 3.8 Flash, high
$0.07
GPT-6 Luna, max effort
Cost per task to run the Intelligence Index v4.3 at list price. Source: Artificial Analysis, September 29, 2026.

At xhigh effort the per-task costs close up, to $3.46 for Opus 5.5 and $2.31 for Astra. List prices fell at OpenAI and Anthropic this month and held at Google:

Slope chart of output price per million tokens, previous to current version: Claude Opus $25 to $20, OpenAI Sol $20 to $10, Claude Sonnet $10 to $10, Gemini Flash $3.75 to $3.75, OpenAI Luna $1.20 to $0.50.

All three vendors take 50% off for batch jobs. I lined up the three price pages, and OpenAI's long-context rule is the one that changes budgets: past 272K input tokens, OpenAI reprices the whole request. Neither Claude's nor Google's price page lists a long-context surcharge for these models. For the older GPT line and its prices, see every OpenAI model compared.

Where the Cheaper Tiers Fit

Claude Sonnet 5.5 scores within 1.6 points of Opus 5.5 on the Artificial Analysis index at half the token price. GPT-6 Sol does the same job for OpenAI at an identical $2 and $10, but scores 47.5.

  • GPT-6 Sol ($2 / $10): OpenAI reports 68.8% on DeepSWE v1.1 and 60.5% on offline OSWorld 2.0; on factual errors flagged by users, it makes about half as many mistakes as GPT-5.6 Sol.
  • GPT-6 Luna ($0.10 / $0.50): the cheapest model here, and free ChatGPT users get it. It scores 83.3 on long-context retrieval (AA-LCR), above Astra.
  • Claude Sonnet 5.5 ($2 / $10): Anthropic reports 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5, and it is the first Sonnet to ship with cyber safeguards.
  • Claude Haiku 4.5 ($1 / $5): still the current Haiku, with a 200K context; Haiku 5.5 has no date yet.
  • Gemini 3.7 Flash and 3.5 Flash-Lite: 3.7 Flash costs the same as 3.8 Flash and uses fewer tokens; Flash-Lite starts at $0.30 per million input tokens.

For open-weight options below these prices, see our MiMo vs Kimi vs Qwen comparison.

Safety Disclosures and Incidents

OpenAI says GPT-6 Astra meets the Critical cybersecurity threshold in its Preparedness Framework, the top tier, a finding it first reached on August 7. In its pacing update it describes a two-week pause in reinforcement learning after its agents broke out of an evaluation and compromised Hugging Face's servers in July.

OpenAI traced that incident to an internal research model, not Astra. OpenAI's incident page says it has notified "dozens of third parties" where agents in training or evaluation bypassed security controls. TechCrunch and Al Jazeera covered the launches against that backdrop.

OpenAI on GPT-6 Astra
  • 0% of cases went beyond the authorised target in its impossible-task test, against 48% for GPT-5.6 Sol without safeguards
  • Written reasoning is harder to monitor than GPT-5.6 Sol's
  • Production classifiers can pause or stop API tasks
Anthropic on Claude Opus 5.5
  • 1.5% of sandbox scenarios drew a containment-crossing attempt, all low severity
  • The model often suspects it is being tested, which limits what audits show
  • METR judged it unlikely to fully automate AI research

I went through the Opus 5.5 system card, which runs past 200 pages. Google's 3.8 Flash model card is a few screens long and points to the 3.7 Flash card for most sections. It does flag one regression: multilingual safety scored 5.4 points worse than Gemini 3.7 Flash.

Which One to Pick by Use Case

If you needPickWhy
Coding agentsClaude Opus 5.5, or Sonnet 5.5 on a budgetFirst in LMArena's code arena; Sonnet 5.5 has the top Terminal-Bench 4.0 run on Artificial Analysis at 63.6%
Computer use and browser agentsGPT-6 Astra72.6% on offline OSWorld 2.0, 92.7% on ScreenSpot-Pro (OpenAI); Anthropic's 81.8% for Opus 5.5 uses a different OSWorld set-up
Long documentsClaude Opus 5.5Top AA-LCR score at 84.7; GPT-6 Luna reaches 83.3 for a fraction of the price under 272K tokens
Cost-sensitive volumeGPT-6 Luna$0.10 and $0.50 per million, $0.07 per index task
Video and audio inputGemini 3.8 FlashThe only one of the three that reads video and audio, at 239 tokens a second
Maths and science researchGPT-6 Astra98% on FrontierMath Tier 4 and the top FrontierMath Erdős score in Epoch's runs

For coding tools, our AI coding agents roundup compares the tools built on these models, and Gemini vs Claude for coding covers the earlier generation. ChatGPT's reach is in our ChatGPT statistics.

Hire Engineers Who Build on These Models

Routing between Opus, Astra and Flash, or keeping an agent under OpenAI's 272K price step, is engineering work. We place vetted LLM developers from Asia, and our LLM developer cost guide shows what a hire costs. Tell us what you are building to get a shortlist.

Frequently Asked Questions

Is there a Gemini 3.8 Pro?

No. Google's 3.8 releases are Flash, Flash Cyber, Live and text-to-speech models. Its newest Pro model on the API price list is Gemini 3.1 Pro Preview.

What is GPT-6 Astra Pro?

A version of Astra that OpenAI includes in ChatGPT Pro, Business and Enterprise plans. The launch post names it but gives no separate benchmarks or API price.

Can I turn off thinking on these models?

Not on Opus 5.5, which Anthropic offers only with thinking enabled. Astra's effort setting runs from low to max, and Gemini 3.8 Flash offers low, medium and high.

Hire LLM engineers.

Pre-vetted senior engineers from Asia at 50 to 70% below US hiring costs, with first profiles in 24 hours.

Hire LLM engineers Apply as talent →
Matt Li

Written by

Matt Li is a tech-driven entrepreneur with deep expertise in global talent strategy, digital experience optimization, e-commerce, and Web3 innovation. He is the Co-Founder of Second Talent, a US-based company that connects businesses with top-tier tech professionals worldwide. Since launching the company in 2024, Matt has led its growth by leveraging technology to streamline remote hiring and scale distributed teams. With a background spanning product, operations, and innovation, Matt brings a cross-disciplinary perspective to the evolving digital economy. His work sits at the intersection of global talent, emerging technology, and scalable digital transformation.

More posts by Matt Li →

Loading available times…