Reddit
C5R's SciUniverse runs models in a real wet lab, and Fable 5.1 leads GPT-6 Astra 45.3% to 32.5%
SciUniverse Level 1 (published Sep 24) has 92 tasks in 17 families across chemistry, biology and materials. Models control instruments and instruct human operators at C5R's Facility-0, for example synthesizing N-benzyl-4-methylbenzamide and confirming it by LC-MS. Pass@1 scores: Claude Fable 5.1 xhigh 45.3% at $40.61 per attempt, GPT-6 Astra 32.5% at $52.37, Opus 5 30.5%, Grok 4.6 26.2%, Gemini 3.8 Flash 14.6%, GPT-5.6 Sol 9.4%. The two r/singularity videos (382 and 152 upvotes) showed only Astra running the lab, but the leaderboard puts Fable ahead at lower cost.
↳ Follow the thread