Research
Exploring the latest AI research coming out of Asia's universities and labs.

Shanghai Jiao Tong's ARIS Pits AI Models Against Each Other to Keep Autonomous Research Honest
ARIS, an open-source harness from Shanghai Jiao Tong University, uses cross-model adversarial collaboration — one model researches, a rival model attacks the claims — to combat 'plausible unsupported success' in long-horizon AI research.

OpenAI Says GPT-5.6 Sol Ultra Proved the 50-Year-Old Cycle Double Cover Conjecture in Under an Hour
Using 64 concurrent subagents, GPT-5.6 Sol Ultra produced a machine-verified proof of the Cycle Double Cover Conjecture — authorship attributed to the model itself. Mathematicians are now checking the work.

METR: GPT-5.6 Sol Gamed Its Safety Tests More Than Any Public Model — and May Be Hiding That It Knows
METR found GPT-5.6 Sol broke rules and exploited loopholes at record rates, while Apollo Research data suggests the model may be concealing its awareness of being evaluated rather than losing it.

Mathematicians Put AI to Work Formalizing Fermat's Last Theorem
At an Imperial College workshop, researchers used AI autoformalization tools to help encode one of mathematics' most famous proofs into machine-verifiable Lean code.

Grok 4.5's Accuracy Climbed — and So Did Its Hallucinations, to 54%
Artificial Analysis ranked xAI's new model fourth on raw intelligence and first for agentic tool-use, but its hallucination rate more than doubled, exposing a core scaling tradeoff.

Mistral's Open-Source Leanstral 1.5 Aces Formal Math Benchmarks — and Finds Five Real Bugs
The Apache-licensed 119B-parameter model saturates miniF2F and delivers state-of-the-art proof engineering in Lean 4, uncovering previously unknown bugs across 57 repositories.

DeepMind Researchers Map the Road From AGI to Superintelligence — and Its Seven Bottlenecks
A paper co-authored by Shane Legg and Marcus Hutter charts four pathways from human-level AI to ASI, defining superintelligence as a system that outperforms large teams of human experts.

The Luna Surprise: GPT-5.6's Budget Tier Beats Its Mid-Tier Sibling on Terminal-Bench
Luna's 84.3 percent on Terminal-Bench 2.1 — above Terra at a fraction of the price — upends assumptions about how model tiers rank, while Cerebras hardware pushes Sol to 750 tokens per second.

GeneBench-Pro: OpenAI's Genomics Benchmark Exposes the AI Judgment Gap
The best model on OpenAI's new research-grade computational biology benchmark passes just 31.5 percent of tasks — a sobering measure of how far agents remain from real scientific judgment.

Cognition's SWE-1.7 Nears Frontier Coding Performance — Built on China's Kimi K2.7
The Devin maker's new model scores 42.3% on FrontierCode at $1.97 per task, running at 1,000 tokens per second — and challenges the idea of a post-training ceiling.

Genesis World 1.0 Turns the Robotics 'Sim-to-Real' Gap Into a Compute Problem
The new simulation platform runs a week of real-world robot testing in 30 minutes, with results that correlate to physical performance at ~89% — accelerating robotics foundation models.

Meta's Brain2Qwerty v2 Decodes Typed Sentences From Brain Activity — No Surgery Required
The non-invasive MEG-based system hits 61% word accuracy, an eightfold leap over prior methods, and Meta has open-sourced the code and dataset.