TL;DR: Claude Opus 5.5 leads the independent Artificial Analysis Intelligence Index at 57.6, ahead of GPT-6 Astra at 52.7 and Gemini 3.8 Flash at 40.9, and its API costs $4 in and $20 out per million tokens against Astra's $10 and $50. GPT-6 Astra is the pick for computer use and hard maths, where OpenAI's computer-use results and Epoch AI's independent maths runs put it first. Gemini 3.8 Flash is the fast, cheap option at $0.75 and $3.75, because Google has not released a Gemini 3.8 Pro.
On ARC-AGI-3, GPT-6 Astra solved 62.7% of the puzzles through the ARC Prize Foundation's standard harness and 99.9% through the adapter OpenAI supplied. OpenAI's launch post quotes the second number.
Anthropic released Claude Opus 5.5 on September 22, 90 minutes before OpenAI's cheaper GPT-6 Sol and Luna, and Google's newest model this month is still a Flash. I read all three launch posts, the Opus 5.5 system card and the price sheets; the scores below come from published evals, not my own runs.
- 1Anthropic reports 66.4% for Opus 5.5 on Terminal-Bench 4.0. Artificial Analysis measured 59.6%, level with Astra's 59.1%.
- 2In LMArena's blind text vote, Opus 5.5 ranks 1st, Gemini 3.8 Flash 10th and GPT-6 Astra 26th.
- 3Every GPT-6 prompt over 272K tokens is billed at twice the input rate and 1.5 times the output rate.
- 4Gemini 3.8 Flash's price doubles to $1.50 and $7.50 on January 1, 2027.
Which Model Is Best Right Now?
Claude Opus 5.5, on the two independent boards that have scored all three. The Artificial Analysis leaderboard put it first on September 29 with 57.6, and it leads four of the five evals below. Its cheaper sibling, Claude Sonnet 5.5, sits second at 56.0 and has the best Terminal-Bench 4.0 run of any model tested.

GPT-6 Astra scores 52.7 and Gemini 3.8 Flash 40.9. Astra's gap is widest on SciCode and Humanity's Last Exam; on Terminal-Bench 4.0 it trails Opus 5.5 by half a point. LMArena's text leaderboard agrees on the order: Opus 5.5 rates 1,509, Gemini 3.8 Flash 1,492 and Astra 1,478, though Opus has only 2,307 votes so far.
Release Dates, Prices and Context Windows
All three flagships read about a million tokens, but only Gemini 3.8 Flash takes audio and video in. Astra and Opus 5.5 accept text and images and return text.
| Model | Released | Input / output / cache read, per 1M | Context / max output | Consumer plans |
|---|---|---|---|---|
| GPT-6 Astra | Sep 3, 2026 | $10 / $50 / $1 | 1,050,000 / 128K | ChatGPT Plus, Pro, Business, Enterprise |
| Claude Opus 5.5 | Sep 22, 2026 | $4 / $20 / $0.20 | 1M / 128K | Claude Pro, Max, Team, Enterprise |
| Gemini 3.8 Flash | Sep 2, 2026 | $0.75 / $3.75 / $0.075 | 1,048,576 / 65,536 | Google AI Pro and Ultra |
| GPT-6 Sol | Sep 22, 2026 | $2 / $10 / $0.20 | 1,050,000 / 128K | ChatGPT Work and Codex, most paid plans |
| GPT-6 Luna | Sep 22, 2026 | $0.10 / $0.50 / $0.01 | 1,050,000 / 128K | Also Free and Go |
| Claude Sonnet 5.5 | Sep 28, 2026 | $2 / $10 / $0.20 | 1M / 128K | All Claude plans, including Free |
Prices come from the OpenAI API pricing page, Claude's pricing page and Google's Gemini API pricing, read on September 29. The six releases landed inside four weeks:
GPT-6 Astra: Computer Use and Maths

GPT-6 Astra is for agents that drive a desktop or browser, and for maths and science work. On OpenAI's launch table it scores 72.6% on the offline OSWorld 2.0 set in about 40 minutes per task, against 65.7% in 75 minutes for GPT-5.6 Sol.
The maths lead has outside backing. Epoch AI scores it 98% on FrontierMath Tier 4 and calls that tier saturated, and on its new FrontierMath Erdős set of 68 unsolved problems Astra scores 3%, where Claude Fable 5.1 scores 0%. Epoch had not published an index score for Opus 5.5 by September 29.
ARC-AGI-3 needs a footnote. The ARC Prize Foundation verified both results below, and the gap comes from what the harness lets the model carry between turns. OpenAI's adapter keeps its opaque reasoning state and compacts long conversations; the standard harness allows visible notes only.

"GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness."
ARC Prize Foundation, September 3, 2026
- Best for: computer-use agents, research maths and science, long-document retrieval inside OpenAI's own tests (96.3% on 8-needle MRCR at 512K to 1M tokens).
- Where it struggles: price and chat preference. It is 26th on LMArena text, and at max effort Artificial Analysis waits a median 291 seconds for the first answer token.
- Price: $10 in, $50 out, $1 cached; Fast mode doubles both the speed and the price.
Claude Opus 5.5: Coding and Knowledge Work at a Lower Price

Claude Opus 5.5 is for coding agents, long migrations and document-heavy office work. Anthropic's launch post says it costs 40% less to run than Opus 5 on typical workloads: tokens are 20% cheaper, cache reads 60% cheaper, and it spends fewer tokens per task.
The Opus 5.5 system card reports 89.9% on SWE-bench Pro, 93.9% on SWE-bench Multilingual and 81.8% partial credit on OSWorld 2.0. OpenAI published no SWE-bench Pro score for Astra. On five evals where both labs do publish a number, Anthropic's own table shows a close race:

Astra wins Terminal-Bench Science and AutomationBench; Opus 5.5 wins Humanity's Last Exam with tools by 10.5 points. The independent LMArena code arena rates Opus 5.5 at 1,827 and Astra at 1,792, with Gemini 3.8 Flash 28th at 1,580.
Anthropic ran its benchmarks with production safeguards on. If a cyber task trips them, Opus 4.8 answers instead, so those scores likely understate the raw model. Our Claude Fable vs Opus vs Sonnet guide covers where Fable 5.1 still fits above it.
- Best for: agentic coding, codebase-wide refactors, finance and legal drafting, chat quality.
- Where it struggles: latency at max effort, where Artificial Analysis clocks a median 681 seconds before the first answer token. It has no thinking-off mode.
- Price: $4 in, $20 out, $0.20 cache reads, $5 cache writes; Fast mode is $8 and $40.
Gemini 3.8 Flash: Google's Only 3.8 Model

Gemini 3.8 Flash is for high-volume agents and anything with video or audio input. Google's newest Pro model is still Gemini 3.1 Pro Preview from February, at $2 and $12, so the "Gemini 3.8" in any comparison this month is a Flash-tier model priced under a fifth of Opus 5.5's input rate.
The Gemini 3.8 Flash model card puts it at 73.7% on DeepSWE v1.1, against 74.0% for Claude Opus 5, and 59.0% on OSWorld 2.0. Google built it on Gemini 3.7 Flash rather than a new base. Its weak spot is Terminal-Bench 4.0 at 19.1%, a figure the public leaderboard and Artificial Analysis both confirm.
"These performance gains stem from a core design choice: 3.8 Flash works harder."
Tulsee Doshi and Raluca Ada Popa, Google, September 2, 2026
Working harder means more thinking tokens, and Google tells efficiency-first users to stay on Gemini 3.7 Flash, which costs the same. Our guide to every Gemini model tracks the Flash and Pro lines, and our Gemini statistics page covers its user numbers.
- Best for: fast, cheap agent loops, video and audio understanding (87.8% on LVBench in Google's agentic run), Google Workspace users.
- Where it struggles: hard terminal tasks, and a price that doubles in January.
- Price: $0.75 in, $3.75 out (thinking tokens included), $0.075 cached.
Where Vendor Scores and Independent Boards Disagree
Anthropic printed 66.4% for Opus 5.5 and 70.6% for Sonnet 5.5 on Terminal-Bench 4.0; Artificial Analysis's own runs gave 59.6% and 63.6%. OpenAI's and Google's claims survive the rerun within a point. Anthropic's own footnote gives Opus 5.5 a standard error of plus or minus 2.6 points, and it scored Opus at xhigh effort.

The official Terminal-Bench leaderboard lists Astra at 58.2% with Codex, but had no Opus 5.5 or Sonnet 5.5 entry on September 29. Harness matters too: the board grades Claude models in Claude Code, Astra in Codex and Gemini in mini-SWE-agent.
The labs also print different scores for the same rival. On OSWorld 2.0, Google's model card lists Claude Opus 5 at 75.4%, while OpenAI's table lists it at 70.2% on the offline subset it reran. SWE-bench Verified cannot settle the coding question: the SWE-bench leaderboard has no entries newer than February 2026. METR tested Opus 5.5 before launch, but it has not published a time horizon for either flagship.
Speed and Time to First Token
Gemini 3.8 Flash answers in about 22 seconds at high effort. Opus 5.5 at max effort thinks for a median 11 minutes before it writes, then streams faster than Astra.

Those waits shrink at lower settings. On the same board, Opus 5.5 at xhigh effort takes a median 136 seconds and Astra at xhigh 124 seconds. Both vendors sell a Fast mode: OpenAI's gives Astra up to twice the speed at twice the price, and Anthropic's gives Opus 5.5 up to 2.5 times the speed at twice the price.
API Price per Million Tokens
Astra's output tokens cost 2.5 times Opus 5.5's, but Astra spends fewer of them. Artificial Analysis ran its full index at each model's top setting, and the bill per task came out lower for Astra:
At xhigh effort the per-task costs close up, to $3.46 for Opus 5.5 and $2.31 for Astra. List prices fell at OpenAI and Anthropic this month and held at Google:

All three vendors take 50% off for batch jobs. I lined up the three price pages, and OpenAI's long-context rule is the one that changes budgets: past 272K input tokens, OpenAI reprices the whole request. Neither Claude's nor Google's price page lists a long-context surcharge for these models. For the older GPT line and its prices, see every OpenAI model compared.
Where the Cheaper Tiers Fit
Claude Sonnet 5.5 scores within 1.6 points of Opus 5.5 on the Artificial Analysis index at half the token price. GPT-6 Sol does the same job for OpenAI at an identical $2 and $10, but scores 47.5.
- GPT-6 Sol ($2 / $10): OpenAI reports 68.8% on DeepSWE v1.1 and 60.5% on offline OSWorld 2.0; on factual errors flagged by users, it makes about half as many mistakes as GPT-5.6 Sol.
- GPT-6 Luna ($0.10 / $0.50): the cheapest model here, and free ChatGPT users get it. It scores 83.3 on long-context retrieval (AA-LCR), above Astra.
- Claude Sonnet 5.5 ($2 / $10): Anthropic reports 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5, and it is the first Sonnet to ship with cyber safeguards.
- Claude Haiku 4.5 ($1 / $5): still the current Haiku, with a 200K context; Haiku 5.5 has no date yet.
- Gemini 3.7 Flash and 3.5 Flash-Lite: 3.7 Flash costs the same as 3.8 Flash and uses fewer tokens; Flash-Lite starts at $0.30 per million input tokens.
For open-weight options below these prices, see our MiMo vs Kimi vs Qwen comparison.
Safety Disclosures and Incidents
OpenAI says GPT-6 Astra meets the Critical cybersecurity threshold in its Preparedness Framework, the top tier, a finding it first reached on August 7. In its pacing update it describes a two-week pause in reinforcement learning after its agents broke out of an evaluation and compromised Hugging Face's servers in July.
OpenAI traced that incident to an internal research model, not Astra. OpenAI's incident page says it has notified "dozens of third parties" where agents in training or evaluation bypassed security controls. TechCrunch and Al Jazeera covered the launches against that backdrop.
- 0% of cases went beyond the authorised target in its impossible-task test, against 48% for GPT-5.6 Sol without safeguards
- Written reasoning is harder to monitor than GPT-5.6 Sol's
- Production classifiers can pause or stop API tasks
- 1.5% of sandbox scenarios drew a containment-crossing attempt, all low severity
- The model often suspects it is being tested, which limits what audits show
- METR judged it unlikely to fully automate AI research
I went through the Opus 5.5 system card, which runs past 200 pages. Google's 3.8 Flash model card is a few screens long and points to the 3.7 Flash card for most sections. It does flag one regression: multilingual safety scored 5.4 points worse than Gemini 3.7 Flash.
Which One to Pick by Use Case
| If you need | Pick | Why |
|---|---|---|
| Coding agents | Claude Opus 5.5, or Sonnet 5.5 on a budget | First in LMArena's code arena; Sonnet 5.5 has the top Terminal-Bench 4.0 run on Artificial Analysis at 63.6% |
| Computer use and browser agents | GPT-6 Astra | 72.6% on offline OSWorld 2.0, 92.7% on ScreenSpot-Pro (OpenAI); Anthropic's 81.8% for Opus 5.5 uses a different OSWorld set-up |
| Long documents | Claude Opus 5.5 | Top AA-LCR score at 84.7; GPT-6 Luna reaches 83.3 for a fraction of the price under 272K tokens |
| Cost-sensitive volume | GPT-6 Luna | $0.10 and $0.50 per million, $0.07 per index task |
| Video and audio input | Gemini 3.8 Flash | The only one of the three that reads video and audio, at 239 tokens a second |
| Maths and science research | GPT-6 Astra | 98% on FrontierMath Tier 4 and the top FrontierMath Erdős score in Epoch's runs |
For coding tools, our AI coding agents roundup compares the tools built on these models, and Gemini vs Claude for coding covers the earlier generation. ChatGPT's reach is in our ChatGPT statistics.
Hire Engineers Who Build on These Models
Routing between Opus, Astra and Flash, or keeping an agent under OpenAI's 272K price step, is engineering work. We place vetted LLM developers from Asia, and our LLM developer cost guide shows what a hire costs. Tell us what you are building to get a shortlist.
Frequently Asked Questions
Is there a Gemini 3.8 Pro?
No. Google's 3.8 releases are Flash, Flash Cyber, Live and text-to-speech models. Its newest Pro model on the API price list is Gemini 3.1 Pro Preview.
What is GPT-6 Astra Pro?
A version of Astra that OpenAI includes in ChatGPT Pro, Business and Enterprise plans. The launch post names it but gives no separate benchmarks or API price.
Can I turn off thinking on these models?
Not on Opus 5.5, which Anthropic offers only with thinking enabled. Astra's effort setting runs from low to max, and Gemini 3.8 Flash offers low, medium and high.
![AI Recruiting Tools. Top 7 AI Recruiting Tools in 2026 [Tried & Tested], by Second Talent.](https://www.secondtalent.com/wp-content/uploads/2026/09/top-ai-recruiting-tools-featured-v2-768x403.jpg)




