Apple Launches 'Apple Upgrade' Lease-to-Own Program With Klarna on July 28Global AI Experts Push Back on US Distillation Claims Against Moonshot's Kimi K3OpenAI Launches Presence, an Enterprise Agent Platform Deployed by Consultants, Not APIsThe Distillation Wars: Why 'Model Theft' Is the New Front in the US-China AI FightAPIs Out, Consultants In: The Enterprise Agent Platform War Has StartedThe Thailand Problem: How Southeast Asia Became the Hole in America's Chip WallAsia Startup Funding Hits Multiyear Peak: $42.8B in Q2, Led by China and AICorgi Reportedly Raises Again at $4B — Its Third Round in Eight WeeksApple Launches 'Apple Upgrade' Lease-to-Own Program With Klarna on July 28Global AI Experts Push Back on US Distillation Claims Against Moonshot's Kimi K3OpenAI Launches Presence, an Enterprise Agent Platform Deployed by Consultants, Not APIsThe Distillation Wars: Why 'Model Theft' Is the New Front in the US-China AI FightAPIs Out, Consultants In: The Enterprise Agent Platform War Has StartedThe Thailand Problem: How Southeast Asia Became the Hole in America's Chip WallAsia Startup Funding Hits Multiyear Peak: $42.8B in Q2, Led by China and AICorgi Reportedly Raises Again at $4B — Its Third Round in Eight WeeksApple Launches 'Apple Upgrade' Lease-to-Own Program With Klarna on July 28Global AI Experts Push Back on US Distillation Claims Against Moonshot's Kimi K3OpenAI Launches Presence, an Enterprise Agent Platform Deployed by Consultants, Not APIsThe Distillation Wars: Why 'Model Theft' Is the New Front in the US-China AI FightAPIs Out, Consultants In: The Enterprise Agent Platform War Has StartedThe Thailand Problem: How Southeast Asia Became the Hole in America's Chip WallAsia Startup Funding Hits Multiyear Peak: $42.8B in Q2, Led by China and AICorgi Reportedly Raises Again at $4B — Its Third Round in Eight Weeks
OpenAI logo
OpenAI / Tech Times
Research

GeneBench-Pro: OpenAI's Genomics Benchmark Exposes the AI Judgment Gap

The best model on OpenAI's new research-grade computational biology benchmark passes just 31.5 percent of tasks — a sobering measure of how far agents remain from real scientific judgment.

D
Daniel ParkAI Correspondent
4 min read

OpenAI has released GeneBench-Pro, a research-level benchmark designed to test whether AI agents can handle the judgment-heavy analysis that real computational biology demands. The results are humbling: even OpenAI's own best configuration, GPT-5.6 Sol Pro, passes just 31.5 percent of tasks at maximum reasoning effort.

What the Benchmark Tests

GeneBench-Pro presents agents with 129 synthetic problems spanning 10 domains and 21 sub-domains — statistical genetics, population genomics, regulatory omics, proteomics, clinical pharmacogenomics, cancer somatic genomics, microbial genomics and forensic genetics among them. Each problem pairs a realistic, deliberately noisy dataset with a target estimand tied to a downstream decision.

Agents work inside an isolated environment with a short prompt, the raw data files, and a standard bioinformatics stack: Python, scientific computing libraries, and domain tools such as PLINK 2.0. There is no hand-holding — exactly the conditions a graduate student or staff scientist would face.

Research Taste, Not Recall

What distinguishes GeneBench-Pro from earlier science benchmarks is its focus on what OpenAI calls "research taste": the chain of judgment calls about which questions a dataset can actually support, when early diagnostics should change the analysis plan, and when a result is decision-ready versus merely computed. These are the skills that separate competent analysis from misleading noise-fitting — and they resist the pattern-matching that saturates conventional benchmarks.

The Scoreboard

At maximum reasoning level, GPT-5.6 Sol Pro scored 31.5 percent and standard Sol reached 28.7 percent. The strongest non-OpenAI model, Anthropic's Claude Opus 4.8, passed 16.0 percent. The Pro-tier configurations — which purchase additional inference-time compute — delivered gains of 2.8 to 7.1 percentage points over their base counterparts, with the budget-tier Luna Pro showing the largest improvement.

The gap between 31.5 percent and the near-saturated scores models post on older science QA benchmarks is the entire point. Frontier models can retrieve biological facts almost perfectly; asked to do biology — messy data, ambiguous signals, consequential decisions — they fail two times out of three.

Why It Matters

Benchmarks shape research incentives, and GeneBench-Pro arrives as labs increasingly pitch AI as a scientific collaborator — from Anthropic's drug-discovery program to Asia's national biomedical AI initiatives in Singapore, Seoul and Shanghai. OpenAI has open-sourced representative tasks, giving competitors a public target. If scores climb the way SWE-bench's did, the milestones will be meaningful ones: each percentage point on GeneBench-Pro represents a genuine unit of scientific judgment, not another memorized fact.

Newsletter

Get Lanceum in your inbox

Weekly insights on AI and technology in Asia.

Share

More in Research

Lanceum

Independent coverage of AI and technology across Asia. We go beyond headlines to explain what matters.

Colophon

Typeset in Space Grotesk & DM Serif Display. Built with Nuxt & Tailwind. Powered by curiosity.

© 2026 Lanceum. All rights reserved.

Independent • Rigorous • Asia-Focused