ExplorationBench Uses 'Alien Worlds' With Rules That Contradict Pretraining to Test Whether Agents Actually Discover Things
arXiv·low signal
arXiv 2609.30199 builds two sandboxes, AlienCode (31 targets, 70 tasks) and AlienLogic (24 targets, 70 tasks), each with executable rules, a deliberately flawed manual and a tool-call schema, so recalled knowledge cannot solve the tasks. Across 10 systems, the strongest learned and applied unfamiliar rules. Results varied widely between trajectories, and continued exploration sometimes stalled or erased earlier gains.