Asia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M Users
Claude Opus 5 announcement artwork from Anthropic
Anthropic
Research

Inside the Claude Opus 5 Numbers: Frontier-Bench Doubled, ARC-AGI-3 at Nearly Four Times GPT-5.6

Anthropic's July 24 release posts a full generational jump over Opus 4.8 — 79.2% on SWE-bench Pro, 43.3% on Frontier-Bench, a GDPval Elo lead — while undercutting the company's own Fable 5 on price. A benchmark-by-benchmark breakdown.

M
Maya SantosSenior Reporter
6 min read

Anthropic shipped Claude Opus 5 on July 24, and the launch-day noise has mostly been about positioning. The benchmark tables deserve a closer read — because the pattern in the numbers says more about where model progress is coming from than any single headline score.

The short version: Opus 5 posts a full generation of gains over Opus 4.8 at the same price, takes outright leads on the hardest new evaluations, and lands within a point of Anthropic's own flagship-class models on coding — at a fraction of the cost per task.

Coding: The Saturation Problem, Then the Real Test

On SWE-bench Verified, the industry's long-standing agentic coding standard, Opus 5 scores 96.0% — a number that essentially retires the benchmark. The action has moved to Scale AI's harder SWE-bench Pro, built by its SEAL evaluation lab specifically because Verified stopped discriminating between frontier models.

There, Opus 5 lands at 79.2%, a ten-point generational jump over Opus 4.8's 69.2%. Notably, the two models ahead of it on the leaderboard are both Anthropic's own: Claude Mythos 5 at 80.3% — the restricted-access, unsafeguarded configuration — and Claude Fable 5, the production-safeguarded version of the same Mythos-class weights, at 80.0%. The top three slots on the industry's toughest coding benchmark now all belong to one lab, separated by 1.1 points.

On CursorBench 3.2, which measures real-world editor-integrated coding, independent trackers place Opus 5 within half a point of Fable 5 — at roughly half the cost.

The New Benchmarks Are the Story

The more interesting results come from evaluations designed after the last generation saturated the old ones. On Frontier-Bench v0.1, Opus 5 scores 43.3% — roughly double Opus 4.8's performance and clearly ahead of both OpenAI's GPT-5.6 Sol (34.4%) and Fable 5 itself (33.7%).

The starkest gap is on ARC-AGI-3, the abstraction-and-reasoning benchmark built to resist memorization: Opus 5 scores 30.2% against GPT-5.6 Sol's 7.8% — nearly a four-fold margin on the evaluation most closely watched by researchers skeptical of benchmark gaming. And on GDPval-AA v2, which measures economically valuable knowledge work as an Elo rating, Opus 5 tops the leaderboard at 1,861, ahead of Fable 5 (1,747) and GPT-5.6 Sol (1,736).

Agentic evaluations follow the same pattern. Trackers report Opus 5 surpassing Fable 5 on OSWorld 2.0 computer-use tasks at roughly a third of the compute budget, and posting pass rates around 1.5 times its competitors on Zapier's AutomationBench.

The Efficiency Lever

The mechanism behind the cost numbers is an "effort" toggle — low, medium and high settings that dial reasoning compute up for hard problems and down for routine ones, with extended thinking on by default and a 1-million-token context window throughout. At $5 per million input tokens and $25 per million output ($10/$50 for a fast mode at roughly 2.5x speed), Opus 5 launches at Opus 4.8's price while, on several evaluations, beating models that cost multiples more per task.

That is the same playbook China's open-weight labs have been running all year — maximum capability per inference dollar — and it is hard not to read Opus 5 partly as an answer to it. As of July 27, BenchLM ranks Opus 5 the top model overall at 85.88/100, but Moonshot's Kimi K3, whose 2.8-trillion-parameter open weights landed the same day, sits at the top of the open class with aggressive sparsity delivering exactly this kind of cost-capability ratio. The competitive frontier is no longer peak score; it is score per dollar.

Caveats

Two things the tables don't show. First, several of the strongest results — Frontier-Bench, GDPval-AA v2 — are new evaluations with short track records, and early leads on young benchmarks have a history of compressing once rivals tune for them. Second, the launch-week numbers mix Anthropic's own reporting with third-party trackers; fully independent replication, particularly on the agentic suites, will take weeks. But the generational delta over Opus 4.8 is consistent across every evaluation with a comparison point — and a ten-point jump on SWE-bench Pro at unchanged pricing is the kind of result that reprices the market whether or not the newer benchmarks hold.

Newsletter

Get Lanceum in your inbox

Weekly insights on AI and technology in Asia.

Share

More in Research

Lanceum

Independent coverage of AI and technology across Asia. We go beyond headlines to explain what matters.

Colophon

Typeset in Space Grotesk & DM Serif Display. Built with Nuxt & Tailwind. Powered by curiosity.

© 2026 Lanceum. All rights reserved.

Independent • Rigorous • Asia-Focused