
Epoch AI: Detectors Rarely Flag Humans — but Miss Up to 48% of AI Text That Mimics Real Authors
New Epoch AI research finds leading detectors like Pangram, GPTZero and Originality.ai almost never falsely accuse human writers, but miss roughly 13% of style-imitated AI passages on average — with miss rates climbing to 48% in scientific writing.
AI text detectors have quietly become infrastructure — deployed by universities, journals, employers and now publishing platforms. New research from Epoch AI delivers a two-sided verdict on whether they deserve the trust: they almost never wrongly flag human writing, but they can be beaten far more often than their marketing suggests, especially where it matters most.
The setup
Epoch tested three leading commercial detectors — Pangram, GPTZero and Originality.ai — against AI-generated passages designed to evade them the way real users do: by prompting models to imitate the style of specific human authors, using five samples of each author's writing as a guide.
The findings
The good news is on the false-positive side: detectors rarely flagged genuine human writing as AI, supporting vendor claims like Pangram's advertised 0.01 percent false-positive rate. For anyone worried about students or journalists being wrongly accused by an algorithm, that is a meaningful result.
The false negatives are the problem. Across 297 style-imitated passages, an average of roughly 13 percent slipped past all three detectors undetected. Individually, Pangram missed 10 percent of style-imitated texts, GPTZero 11 percent, and Originality.ai 18 percent.
Genre changed everything. For scientific writing, miss rates climbed as high as 48 percent — Pangram failed to catch 25 percent of style-imitated academic AI text, GPTZero 24 percent, and Originality.ai 29 percent. The genre where detection is deployed most aggressively — journal submissions, peer review, student research — is precisely where it fails most.
The structural asymmetry
The result illustrates a defensive dilemma the paper's readers will recognize from security research: making a model imitate a writing style takes one sentence of prompting, while detecting the result requires distinguishing subtle statistical patterns from an ever-improving generator. The attacker's costs fall with every model release; the defender's rise.
Why it lands hard this week
The study arrived days before Substack launched Pangram-powered AI detection for its readers — with framing that suddenly looks well-advised: detection surfaced as an "estimate" and reader context, never as automatic enforcement. Epoch's numbers suggest that is the only defensible deployment. As institutions across Asia and the West wire detectors into consequential decisions — admissions, publication, hiring — the research draws a clear line: current tools are good enough to inform human judgment, and nowhere near good enough to replace it.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.


