
GeneBench-Pro: OpenAI's Genomics Benchmark Exposes the AI Judgment Gap
The best model on OpenAI's new research-grade computational biology benchmark passes just 31.5 percent of tasks — a sobering measure of how far agents remain from real scientific judgment.
OpenAI has released GeneBench-Pro, a research-level benchmark designed to test whether AI agents can handle the judgment-heavy analysis that real computational biology demands. The results are humbling: even OpenAI's own best configuration, GPT-5.6 Sol Pro, passes just 31.5 percent of tasks at maximum reasoning effort.
What the Benchmark Tests
GeneBench-Pro presents agents with 129 synthetic problems spanning 10 domains and 21 sub-domains — statistical genetics, population genomics, regulatory omics, proteomics, clinical pharmacogenomics, cancer somatic genomics, microbial genomics and forensic genetics among them. Each problem pairs a realistic, deliberately noisy dataset with a target estimand tied to a downstream decision.
Agents work inside an isolated environment with a short prompt, the raw data files, and a standard bioinformatics stack: Python, scientific computing libraries, and domain tools such as PLINK 2.0. There is no hand-holding — exactly the conditions a graduate student or staff scientist would face.
Research Taste, Not Recall
What distinguishes GeneBench-Pro from earlier science benchmarks is its focus on what OpenAI calls "research taste": the chain of judgment calls about which questions a dataset can actually support, when early diagnostics should change the analysis plan, and when a result is decision-ready versus merely computed. These are the skills that separate competent analysis from misleading noise-fitting — and they resist the pattern-matching that saturates conventional benchmarks.
The Scoreboard
At maximum reasoning level, GPT-5.6 Sol Pro scored 31.5 percent and standard Sol reached 28.7 percent. The strongest non-OpenAI model, Anthropic's Claude Opus 4.8, passed 16.0 percent. The Pro-tier configurations — which purchase additional inference-time compute — delivered gains of 2.8 to 7.1 percentage points over their base counterparts, with the budget-tier Luna Pro showing the largest improvement.
The gap between 31.5 percent and the near-saturated scores models post on older science QA benchmarks is the entire point. Frontier models can retrieve biological facts almost perfectly; asked to do biology — messy data, ambiguous signals, consequential decisions — they fail two times out of three.
Why It Matters
Benchmarks shape research incentives, and GeneBench-Pro arrives as labs increasingly pitch AI as a scientific collaborator — from Anthropic's drug-discovery program to Asia's national biomedical AI initiatives in Singapore, Seoul and Shanghai. OpenAI has open-sourced representative tasks, giving competitors a public target. If scores climb the way SWE-bench's did, the milestones will be meaningful ones: each percentage point on GeneBench-Pro represents a genuine unit of scientific judgment, not another memorized fact.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.


