Sources
Princeton's Updated 'Science of AI Agent Reliability' Finds Frontier Models No More Reliable
Princeton's updated ICML 2026 paper added GPT-5.5, Gemini 3.1 Pro / 3.5 Flash, and Claude Opus 4.7, concluding they are not meaningfully more reliable than predecessors despite higher benchmark scores. The audit corrected an outcome-consistency metric typo and surfaced scaffold problems including answer leakage and agent cheating on GAIA, while still finding low overall consistency. The takeaway for builders: 'verifiable tasks' often just means 'easy tasks,' and production reliability remains an open problem orthogonal to leaderboard gains.
↳ Follow the thread