
Grok 4.5's Accuracy Climbed — and So Did Its Hallucinations, to 54%
Artificial Analysis ranked xAI's new model fourth on raw intelligence and first for agentic tool-use, but its hallucination rate more than doubled, exposing a core scaling tradeoff.
xAI's Grok 4.5, launched into a crowded field alongside GPT-5.6 and Anthropic's latest, delivered exactly the kind of split verdict that has become the signature of frontier model releases: measurably smarter on aggregate benchmarks, best-in-class at agentic tool use — and dramatically more prone to confidently stating things that aren't true.
The Scores
Artificial Analysis placed Grok 4.5 fourth on its Intelligence Index with a score of 54, trailing only Fable 5, GPT-5.5, and Claude Opus 4.8. On agentic tool-use — the ability to plan and execute multi-step tasks with external tools — it ranked first, a genuinely notable result given how central agentic capability has become to real-world deployment. Snorkel's GDPval+ evaluation showed Grok 4.5 posting a 29% mean pass rate against competitors at 21–22%.
By the numbers that dominate launch-day headlines, in other words, Grok 4.5 is a strong model, and Elon Musk's "Opus-class" framing isn't baseless.
The Catch
Then there's the AA-Omniscience Index, which measures factual accuracy and, crucially, hallucination — how often a model confidently asserts false information rather than declining to answer. Here Grok 4.5's numbers tell a darker story: accuracy improved from 35% to 52%, but the hallucination rate jumped from 25% on the prior generation to 54%.
Read carefully, that means on the knowledge tasks where the model is wrong, it now expresses confidence in its incorrect answer more than half the time. The model knows more — and is more dangerously certain when it doesn't.
A Pattern, Not an Anomaly
Artificial Analysis framed this as a recurring dynamic: larger, more capable models tend to "know more but are also more confident in their knowledge." As models scale, they absorb more facts, which raises accuracy — but they also become less calibrated, less willing to say "I don't know," which raises hallucination. Capability and calibration are pulling in opposite directions.
This is the tradeoff that benchmark leaderboards systematically obscure. A single "intelligence" number rewards breadth of knowledge and reasoning while saying nothing about whether the model knows the limits of what it knows. For a chatbot answering trivia, a 54% hallucination rate on hard questions is an annoyance. For an agentic system taking autonomous actions — the very use case where Grok 4.5 excels — a confidently wrong step can cascade into real consequences.
What Buyers Should Take Away
The Grok 4.5 results are a useful corrective to leaderboard tunnel vision. The right model for a task depends less on its rank on an aggregate index than on the shape of its errors: a coding agent needs reliability under tool use; a research assistant needs calibrated uncertainty; a customer-facing bot needs to refuse gracefully rather than fabricate.
As the frontier compresses — five near-peer models now launch within weeks of one another — differentiation is shifting from raw capability to these reliability characteristics. Grok 4.5's strong agentic scores make it a serious contender for automation workloads. Its hallucination numbers are a reminder that "smarter" and "more trustworthy" are not the same axis, and that in 2026, buyers need to measure both.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.


