Fetching from the wire…
Research2026-08-20 · source-backed
A Princeton-led study with Stanford, Berkeley, Johns Hopkins, Toronto, Georgetown and the UK AI Security Institute handed Claude Opus 4.8 and GPT-5.6 Sol Ultra the central research question from unpublished NeurIPS 2026 submissions. The original authors graded the output as reviewers and rejected both, citing weak experimental design, unsupported conclusions and impenetrable prose. Neither agent spent its full budget. (MIT Technology Review) Kirgis and Kapoor argue the failure was judgment, not engineering: the agents could run experiments and write LaTeX, they just couldn't decide which hypotheses deserved compute.
Each link below shares sources, entities, or timing with this story.
Stanford benchmarked against Claude / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Stanford benchmarked against Claude); both cover MIT Technology Review, Stanford; reported by the same outlet (technologyreview.com).
Stanford benchmarked against Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against Claude); both cover Claude Opus, GPT; overlapping topics (agent, claude).
Linked by a graph relationship (Stanford benchmarked against Claude); both cover GPT, Sol Ultra; overlapping topics (agent, claude).
Stanford benchmarked against Gemini / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against Gemini); both cover Claude Opus, GPT; overlapping topics (agent, claude).
Stanford benchmarked against Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against Claude); both cover Claude Opus, GPT; overlapping topics (agent, claude).
Stanford benchmarked against DeepSeek / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against DeepSeek); both cover Claude Opus, GPT; overlapping topics (claude, credit).
Stanford benchmarked against Claude / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against Claude); both cover GPT, MIT Technology Review; reported by the same outlet (technologyreview.com).
Stanford benchmarked against Gemini / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against Gemini); both cover Claude Opus, GPT; overlapping topics (agent, claude).