
Can You Prove a Model Was Distilled? Inside the Forensics Behind the Kimi K3 Dispute
Redwood Research's Ryan Greenblatt published cross-entropy analysis showing K3 identifies itself as Claude at anomalous rates — evidence that is suggestive, statistical, and stubbornly short of proof.
As Washington and Beijing trade accusations over whether Moonshot's Kimi K3 was built on Anthropic's Fable, the most interesting work is happening away from the podiums: researchers are testing whether model distillation can be detected at all from the outside.
The cross-entropy fingerprint
Redwood Research chief scientist Ryan Greenblatt published an analysis examining K3's behavioral statistics, including a striking pattern: across many prompts, K3 identifies itself as Claude at rates far above what unrelated models exhibit. Cross-entropy measurements — how "surprised" one model is by another's characteristic outputs — showed statistical patterns consistent with shared training lineage between K3 and Anthropic's model family.
Self-identification quirks are the classic tell of output-trained models. When a model insists it is Claude, the most parsimonious explanation is that Claude transcripts — harvested directly or scraped from the public web — sat somewhere in its training data. The cross-entropy work goes further, quantifying distributional similarity rather than cherry-picking embarrassing responses.
Why it still isn't proof
Greenblatt himself flagged the caveats. The modern web is saturated with Claude-generated text; any model trained on 2025-26 web scrapes ingests Anthropic outputs without a single API call. Contaminated public datasets, shared RLHF vendors and convergent post-training recipes can all produce similar fingerprints. Behavioral forensics can establish that K3's training data contains substantial Claude-derived text — it cannot distinguish deliberate industrial harvesting from the ambient contamination that afflicts every frontier model, in every country.
That gap matters because the policy stakes are calibrated to intent. Anthropic's telemetry — some 16 million fake-account interactions attributed to Chinese labs — speaks to deliberate collection, but connecting specific harvested tokens to specific K3 weights is beyond any published technique. Weights carry no watermark that survives training; there is no cryptographic chain of custody for knowledge.
The research agenda this fight just created
Expect a surge of work on exactly this problem: output watermarking robust to distillation, canary-string protocols for API responses, and statistical provenance tests with known false-positive rates. The field has an uncomfortable deadline — K3's weights go public on July 27, at which point the world's researchers can run every test they like, and both governments will be quoting whichever results suit them.
The deeper lesson from the episode is that the AI industry has built its intellectual-property regime on terms-of-service documents and the honor system, then scaled it to geopolitics. The forensics are improving fast. The proof standard the accusations require may simply not exist.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Kunlun's Mureka V9.5 Targets the 'AI Smell' in Machine-Made Music With Reflective Reasoning

DeepSeek V4 Stable Release Lands July 24 — and the Open Trillion-Parameter Class Gets Crowded
