Asia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M Users
Benchmark chart illustrating Grok 4.5's intelligence and hallucination metrics
Artificial Analysis
Research

Grok 4.5's Accuracy Climbed — and So Did Its Hallucinations, to 54%

Artificial Analysis ranked xAI's new model fourth on raw intelligence and first for agentic tool-use, but its hallucination rate more than doubled, exposing a core scaling tradeoff.

D
Daniel ParkAI Correspondent
4 min read

xAI's Grok 4.5, launched into a crowded field alongside GPT-5.6 and Anthropic's latest, delivered exactly the kind of split verdict that has become the signature of frontier model releases: measurably smarter on aggregate benchmarks, best-in-class at agentic tool use — and dramatically more prone to confidently stating things that aren't true.

The Scores

Artificial Analysis placed Grok 4.5 fourth on its Intelligence Index with a score of 54, trailing only Fable 5, GPT-5.5, and Claude Opus 4.8. On agentic tool-use — the ability to plan and execute multi-step tasks with external tools — it ranked first, a genuinely notable result given how central agentic capability has become to real-world deployment. Snorkel's GDPval+ evaluation showed Grok 4.5 posting a 29% mean pass rate against competitors at 21–22%.

By the numbers that dominate launch-day headlines, in other words, Grok 4.5 is a strong model, and Elon Musk's "Opus-class" framing isn't baseless.

The Catch

Then there's the AA-Omniscience Index, which measures factual accuracy and, crucially, hallucination — how often a model confidently asserts false information rather than declining to answer. Here Grok 4.5's numbers tell a darker story: accuracy improved from 35% to 52%, but the hallucination rate jumped from 25% on the prior generation to 54%.

Read carefully, that means on the knowledge tasks where the model is wrong, it now expresses confidence in its incorrect answer more than half the time. The model knows more — and is more dangerously certain when it doesn't.

A Pattern, Not an Anomaly

Artificial Analysis framed this as a recurring dynamic: larger, more capable models tend to "know more but are also more confident in their knowledge." As models scale, they absorb more facts, which raises accuracy — but they also become less calibrated, less willing to say "I don't know," which raises hallucination. Capability and calibration are pulling in opposite directions.

This is the tradeoff that benchmark leaderboards systematically obscure. A single "intelligence" number rewards breadth of knowledge and reasoning while saying nothing about whether the model knows the limits of what it knows. For a chatbot answering trivia, a 54% hallucination rate on hard questions is an annoyance. For an agentic system taking autonomous actions — the very use case where Grok 4.5 excels — a confidently wrong step can cascade into real consequences.

What Buyers Should Take Away

The Grok 4.5 results are a useful corrective to leaderboard tunnel vision. The right model for a task depends less on its rank on an aggregate index than on the shape of its errors: a coding agent needs reliability under tool use; a research assistant needs calibrated uncertainty; a customer-facing bot needs to refuse gracefully rather than fabricate.

As the frontier compresses — five near-peer models now launch within weeks of one another — differentiation is shifting from raw capability to these reliability characteristics. Grok 4.5's strong agentic scores make it a serious contender for automation workloads. Its hallucination numbers are a reminder that "smarter" and "more trustworthy" are not the same axis, and that in 2026, buyers need to measure both.

Newsletter

Get Lanceum in your inbox

Weekly insights on AI and technology in Asia.

Share

More in Research

Lanceum

Independent coverage of AI and technology across Asia. We go beyond headlines to explain what matters.

Colophon

Typeset in Space Grotesk & DM Serif Display. Built with Nuxt & Tailwind. Powered by curiosity.

© 2026 Lanceum. All rights reserved.

Independent • Rigorous • Asia-Focused