Research
Exploring the latest AI research coming out of Asia's universities and labs.

Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo
Moonshot's release of the largest open-weight model ever — under a near-permissive license, with day-zero third-party hosting — resets the open/closed gap and hands every enterprise a frontier-class model it can run itself.

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake
A new multimodal benchmark tests whether AI can detect 'pseudo-consensus' in meetings — apparent agreement masking private dissent — and finds today's models still can't read the room.

SciCodePile: A 128GB Corpus and a Brutal Benchmark Where Top Models Solve 12% of Scientific Code Tasks
A new executable benchmark drawn from 37,737 research repositories exposes a gap the coding-agent hype has papered over: on real scientific code, frontier models pass barely one task in eight.

Inside the Claude Opus 5 Numbers: Frontier-Bench Doubled, ARC-AGI-3 at Nearly Four Times GPT-5.6
Anthropic's July 24 release posts a full generational jump over Opus 4.8 — 79.2% on SWE-bench Pro, 43.3% on Frontier-Bench, a GDPval Elo lead — while undercutting the company's own Fable 5 on price. A benchmark-by-benchmark breakdown.

FLUX 3: Black Forest Labs Trains One Model to Generate Video, Audio and Robot Actions
The German lab's new multimodal flow model produces 20-second clips with natively synchronized audio from a single set of weights — and extends the same architecture to robotic action prediction, already tested on Audi production lines.

NVIDIA and KAIST Launch $300 Million Agentic AI Lab in Seoul
A five-year joint research laboratory at KAIST's Kim Jaechul Graduate School of AI will build Korea-optimized agentic models on NVIDIA's Nemotron stack, backed by $50 million a year in compute — the latest move in South Korea's sovereign AI push.

Kunlun's Mureka V9.5 Targets the 'AI Smell' in Machine-Made Music With Reflective Reasoning
Unveiled at WAIC alongside the O3 music reasoning model and Matrix-Game 3.5, the new system trades maximalist arrangements for restraint — and an agentic creation loop that critiques its own compositions.

Can You Prove a Model Was Distilled? Inside the Forensics Behind the Kimi K3 Dispute
Redwood Research's Ryan Greenblatt published cross-entropy analysis showing K3 identifies itself as Claude at anomalous rates — evidence that is suggestive, statistical, and stubbornly short of proof.

DeepSeek V4 Stable Release Lands July 24 — and the Open Trillion-Parameter Class Gets Crowded
DeepSeek's V4 stable build arrives this week, squaring off against Kimi K3 and GLM-5.2 in a three-way fight among open trillion-scale MoE models — with benchmark parity against closed frontiers at a fraction of the serving cost.

Epoch AI: Detectors Rarely Flag Humans — but Miss Up to 48% of AI Text That Mimics Real Authors
New Epoch AI research finds leading detectors like Pangram, GPTZero and Originality.ai almost never falsely accuse human writers, but miss roughly 13% of style-imitated AI passages on average — with miss rates climbing to 48% in scientific writing.

Google Begins Pretraining Gemini 4, Calling It Its 'Most Ambitious Run Yet' — With 3.5 Pro Still Missing
Google confirmed Gemini 4 has entered pretraining as a completely revamped foundation model, after the original Gemini 3.5 Pro base was reportedly scrapped and restarted — with the Frozen v2 chip program hovering in the background.

Benchmarking the Open-Weight War: Chinese Models Now Match the Entire US Open Ecosystem
A new independent analysis finds Qwen, Kimi, DeepSeek and GLM dominating the open-weight frontier, with Moonshot's K2 Thinking at Intelligence Index 67 — and America's best open model trailing on knowledge.