Research
Exploring the latest AI research coming out of Asia's universities and labs.

Google Ships Gemini 3.8 Flash and a 'Cyber' Twin That Hunts Vulnerabilities
The new Flash flagship jumps to 90.8% on Terminal-Bench 2.1 and beats larger frontier models on long-horizon coding, while Gemini 3.8 Flash Cyber debuts under Google's restricted Fairwind program for governments.

Meta's Muse Spark 1.3 Reaches the Frontier: 75.4% on DeepSWE, 1M Context, and a Data-for-Discount Endpoint
Meta's strongest model yet posts its biggest jump on coding and agentic benchmarks — and introduces a 'contributor' pricing tier that is 10-20x cheaper if you let Meta train on your data.

Qwen3.8-Flash-Next: Alibaba Previews the Qwen4 Architecture With Hybrid Attention and 'Engram' Embeddings
The experimental open-weights release activates just 6B of 125B parameters per token, pairs Gated DeltaNet with Qwen Sparse Attention, and adds a 51B n-gram lookup table — a blueprint for where China's most-downloaded model family goes next.

Claude Completes First Fully Machine-Verified Proof of Fermat's Last Theorem in 11 Days
Working largely autonomously, Anthropic's model wrote 13 million lines of Lean and proved 29,500 intermediate theorems to formalize a proof mathematicians expected would take years to encode — generating 6 billion tokens across dozens of agents.

HUMAIN's humain-m3: A 428B-Parameter Arabic Frontier Model Trained on a Trillion Arabic Tokens
Saudi Arabia's state-backed lab unveils a mixture-of-experts model built on MiniMax's M3 lineage that tops seven public Arabic benchmarks with an 89.37% average — now in research preview on HUMAIN Node.

Z.ai's GLM-5.3-Flash Pairs Sparse and Linear Attention in an Open-Source First — at $0.15 per Million Tokens
The MIT-licensed 320B model activates just 18B parameters per token, handles a 1M context natively across text and images, and cuts attention compute roughly 3x — undercutting frontier pricing by two orders of magnitude.

Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo
Moonshot's release of the largest open-weight model ever — under a near-permissive license, with day-zero third-party hosting — resets the open/closed gap and hands every enterprise a frontier-class model it can run itself.

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake
A new multimodal benchmark tests whether AI can detect 'pseudo-consensus' in meetings — apparent agreement masking private dissent — and finds today's models still can't read the room.

SciCodePile: A 128GB Corpus and a Brutal Benchmark Where Top Models Solve 12% of Scientific Code Tasks
A new executable benchmark drawn from 37,737 research repositories exposes a gap the coding-agent hype has papered over: on real scientific code, frontier models pass barely one task in eight.

Inside the Claude Opus 5 Numbers: Frontier-Bench Doubled, ARC-AGI-3 at Nearly Four Times GPT-5.6
Anthropic's July 24 release posts a full generational jump over Opus 4.8 — 79.2% on SWE-bench Pro, 43.3% on Frontier-Bench, a GDPval Elo lead — while undercutting the company's own Fable 5 on price. A benchmark-by-benchmark breakdown.

FLUX 3: Black Forest Labs Trains One Model to Generate Video, Audio and Robot Actions
The German lab's new multimodal flow model produces 20-second clips with natively synchronized audio from a single set of weights — and extends the same architecture to robotic action prediction, already tested on Audi production lines.

NVIDIA and KAIST Launch $300 Million Agentic AI Lab in Seoul
A five-year joint research laboratory at KAIST's Kim Jaechul Graduate School of AI will build Korea-optimized agentic models on NVIDIA's Nemotron stack, backed by $50 million a year in compute — the latest move in South Korea's sovereign AI push.