
GeneBench-Pro: OpenAI's Genomics Benchmark Exposes the AI Judgment Gap
The best model on OpenAI's new research-grade computational biology benchmark passes just 31.5 percent of tasks — a sobering measure of how far agents remain from real scientific judgment.
OpenAI has released GeneBench-Pro, a research-level benchmark designed to test whether AI agents can handle the judgment-heavy analysis that real computational biology demands. The results are humbling: even OpenAI's own best configuration, GPT-5.6 Sol Pro, passes just 31.5 percent of tasks at maximum reasoning effort.
What the Benchmark Tests
GeneBench-Pro presents agents with 129 synthetic problems spanning 10 domains and 21 sub-domains — statistical genetics, population genomics, regulatory omics, proteomics, clinical pharmacogenomics, cancer somatic genomics, microbial genomics and forensic genetics among them. Each problem pairs a realistic, deliberately noisy dataset with a target estimand tied to a downstream decision.
Agents work inside an isolated environment with a short prompt, the raw data files, and a standard bioinformatics stack: Python, scientific computing libraries, and domain tools such as PLINK 2.0. There is no hand-holding — exactly the conditions a graduate student or staff scientist would face.
Research Taste, Not Recall
What distinguishes GeneBench-Pro from earlier science benchmarks is its focus on what OpenAI calls "research taste": the chain of judgment calls about which questions a dataset can actually support, when early diagnostics should change the analysis plan, and when a result is decision-ready versus merely computed. These are the skills that separate competent analysis from misleading noise-fitting — and they resist the pattern-matching that saturates conventional benchmarks.
The Scoreboard
At maximum reasoning level, GPT-5.6 Sol Pro scored 31.5 percent and standard Sol reached 28.7 percent. The strongest non-OpenAI model, Anthropic's Claude Opus 4.8, passed 16.0 percent. The Pro-tier configurations — which purchase additional inference-time compute — delivered gains of 2.8 to 7.1 percentage points over their base counterparts, with the budget-tier Luna Pro showing the largest improvement.
The gap between 31.5 percent and the near-saturated scores models post on older science QA benchmarks is the entire point. Frontier models can retrieve biological facts almost perfectly; asked to do biology — messy data, ambiguous signals, consequential decisions — they fail two times out of three.
Why It Matters
Benchmarks shape research incentives, and GeneBench-Pro arrives as labs increasingly pitch AI as a scientific collaborator — from Anthropic's drug-discovery program to Asia's national biomedical AI initiatives in Singapore, Seoul and Shanghai. OpenAI has open-sourced representative tasks, giving competitors a public target. If scores climb the way SWE-bench's did, the milestones will be meaningful ones: each percentage point on GeneBench-Pro represents a genuine unit of scientific judgment, not another memorized fact.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake
