Three frontier labs shipped new flagship models within a nine-month span: OpenAI’s GPT-5.2 on December 11, 2025, Google’s Gemini 3 on November 18, 2025, and Anthropic’s Claude Opus 5 on July 24, 2026. Each company describes its model as the strongest available for professional and agentic work. Stripped of marketing language, the verifiable differences are narrower than the launch posts suggest, and increasingly about price and efficiency rather than raw capability. All three releases arrived inside a compressed rhythm of updates — this is not one launch cycle a year anymore, but a nearly continuous sequence of point releases, each accompanied by a fresh set of benchmark claims that only partly overlap with the ones the previous release used.
GPT-5.2 targets professional-task performance
OpenAI positions GPT-5.2 as “the most advanced frontier model for everyday professional work.” The company’s release notes report a score of 70.9% on GDPval, a benchmark built from real work products across 44 occupations, which OpenAI says is the first time a model has approached expert-level performance on that measure. On coding, GPT-5.2 reaches 80% on SWE-bench Verified and 55.6% on the harder SWE-bench Pro suite, and 98.7% on Tau2-bench Telecom, a test of multi-turn tool-calling. OpenAI also says vision error rates on chart reading and interface interpretation dropped roughly in half versus GPT-5.1.
Access starts at $1.75 per million input tokens and $14 per million output tokens for the standard tier; the GPT-5.2 Pro tier costs $21 and $168 respectively, with a 90% discount available on cached inputs. OpenAI frames the model around “economically valuable” work rather than open-ended chat, a shift that lines up with how GDPval itself is built: instead of trivia or math puzzles, it grades real deliverables — memos, spreadsheets, design files — against work produced by paid professionals in the same occupations, then has other professionals judge which is better.
Claude Opus 5 chases efficiency at the frontier
Anthropic’s framing for Opus 5, released July 24, 2026, is explicitly about cost: the company describes it as coming “close to the frontier intelligence of Claude Fable 5 at half the price,” referring to Anthropic’s largest and most expensive model. On Anthropic’s internal Frontier-Bench v0.1 for software engineering, Opus 5 more than doubles the score of its predecessor, Opus 4.8. It also scores roughly three times higher than the next-best model on ARC-AGI-3, a benchmark built to test novel problem-solving rather than memorized patterns, and beats Fable 5 on the OSWorld 2.0 computer-use benchmark at about a third of the cost, according to Anthropic.
Pricing is unchanged from Opus 4.5: $5 per million input tokens and $25 per million output tokens. A faster variant runs about 2.5 times quicker at double the base price. Anthropic also reports gains on narrower evaluations that matter to enterprise buyers rather than researchers: roughly 1.5 times the pass rate of competing models on Zapier’s AutomationBench, a suite built around real workflow-automation tasks, plus noted improvements on life-sciences evaluations covering organic chemistry and protein-related work, and stronger data visualization output — capabilities aimed squarely at the agentic and scientific-research customers Anthropic has courted since Opus 4.
Gemini 3 leads general leaderboards, then keeps iterating
Gemini 3 launched November 18, 2025, with Gemini 3 Pro in preview and a 1-million-token context window. At launch, Google cited a 1501 Elo score on the crowd-voted LMArena leaderboard, 37.5% on Humanity’s Last Exam, and 91.9% on GPQA Diamond, a graduate-level science benchmark, plus 81% on MMMU-Pro and 87.6% on Video-MMMU for multimodal reasoning. Google also tied the launch to its existing distribution, noting AI Overviews in Search already reach roughly 2 billion users a month, a scale advantage neither OpenAI nor Anthropic can currently claim for a comparable free-tier product. Independent reviewers, including outlets that ran their own comparative tests rather than repeating vendor claims, generally found Gemini 3 competitive with or ahead of GPT-5.1 and Claude Opus 4.5 on reasoning-heavy tasks at launch, though such reviews predate the release of both GPT-5.2 and Opus 5 and so say little about where the three-way race stands today.
Google has iterated quickly since the November launch: by September 2026 its lineup includes Gemini 3.1 Pro Preview and Gemini 3.8 Flash, both released within months of the original Gemini 3. According to Google’s own API pricing page, the 3.1 Pro Preview costs $2 per million input tokens for prompts up to 200,000 tokens (rising to $4 above that) and $12 to $18 per million output tokens depending on prompt length — figures that were not published anywhere in Google’s original Gemini 3 announcement.
The benchmarks are contested
None of these figures should be read as an independent ranking. GDPval, SWE-bench and ARC-AGI are built or maintained partly outside the labs that cite them, but each vendor selectively publishes results, and LMArena’s crowd-voted scores have drawn criticism for rewarding response style over substance and for letting vendors test private model variants before public release. GDPval itself is an OpenAI-built and OpenAI-evaluated benchmark. Crucially, none of the three companies published head-to-head scores on identical third-party test suites at the same point in time, so any cross-vendor comparison — including the ones above — relies on separately reported, non-simultaneous results rather than a single controlled test.
What it actually costs to use them
Per-token pricing has diverged more visibly than capability. Opus 5 is Anthropic’s priciest standard model at $5/$25 per million tokens; GPT-5.2’s standard tier undercuts it on input cost ($1.75) but exceeds it on output cost ($14); Gemini’s newest Pro preview lands roughly in between on output pricing but was Google’s least transparent launch on pricing, since neither the original Gemini 3 nor the 3 Pro preview post disclosed cost figures at release. For most buyers evaluating these systems in September 2026, the practical question has shifted from “which model wins a benchmark” to which vendor’s coding, agent and enterprise tooling fits an existing stack at a workable price.
Sources
- OpenAI — Introducing GPT-5.2
- Anthropic — Introducing Claude Opus 5
- Google — Gemini 3: Introducing the latest Gemini AI model
- Google AI for Developers — Gemini API pricing
- CNBC — Google announces Gemini 3 as battle with OpenAI intensifies

