Reddit
BixBench3 Tops Out at 0.48: GPT-5.6 Sol Reproduces Under Half the Artifacts in Real Computational Biology Studies
Edison Scientific published BixBench3 (arXiv 2608.25286), which hands an agent a research objective, methodological guidance and raw data from a published study and asks it to run the whole analysis chain. Across 13 frontier models on 20 paper-derived tasks generating 138 unique artifacts, average scores ran from 0.00 for Gemini 3.1 Flash Lite to 0.48 for GPT-5.6 Sol, where 0.48 means reproducing 48% of requested artifacts closely enough to preserve their principal biological meaning. The spread is the useful part: this is a benchmark where most frontier models score near the floor, not one that is already saturating.
↳ Follow the thread