Research
AI Scientists Benchmarked Against Formula 1 Designers and Magic Pro Tour Decks: GPT-5.2 Matched 10 of 40 Real Innovations From 166 Ideas
To avoid synthetic tasks and retrospective contamination, this benchmark uses adversarial fast-moving domains where expert practitioners independently produce observable ground truth. In F1, models ideated 2026-season car design concepts against real pre-season innovations; GPT-5.2 matched 10 of 40 across 166 proposed ideas. In Magic: The Gathering, Gemini 3 Flash's best deck recovered 5 of 7 new-set cards from the third-place Pro Tour deck, and across all 108 generated decks the cards models picked most often were also those most adopted by Pro Tour decks (Spearman ρ = 0.74, p = 0.0003) — pointing to filtering and prioritization, not idea generation, as the capability gap.
↳ Follow the thread