JuliaHub Physical AI Eval: Swapping the Agent Harness Moved Scores More Than Twice as Much as Swapping the Frontier Model
JuliaHub / Hacker News·high signal
A JuliaHub study published 2026-07-30 ran four frontier models through five sealed modeling and simulation problems escalating from constitutive consistency to a full NASA HL-20 flight vehicle with six-degree-of-freedom dynamics. Claude Fable 5 took the difficulty-weighted top score at 0.889 -- the only model to sweep all twelve trials on the four core problems -- ahead of GPT-5.6-Sol at 0.814, GPT-5.6-Terra at 0.786 and GPT-5.6-Luna at 0.727. The buried headline: changing the agent harness produced a 0.366-point spread, more than double the 0.162 gap between the best and worst models.