OpenAI Introduces GeneBench-Pro to Benchmark AI Agents in Computational Biology
OpenAI Blog·medium signal
OpenAI released GeneBench-Pro, a research-level benchmark for judging AI agents on genomics, biology, and scientific-research tasks using harder, more realistic synthetic datasets that extend the original GeneBench. It arrives the same week as Anthropic's Claude Science, underscoring a competitive push toward AI-for-science evaluation and tooling. For builders, it's a new yardstick for agent performance on real computational-biology workflows.