Sources
TaxiGPT does have a world model; its failures come from feature interference, not a missing map
Submitted 2026-09-18, 2609.21748 revisits a transformer trained on random walks through Manhattan whose behavioral failures were read as evidence of an incoherent internal map. Mechanistic analysis and causal interventions show the model represents intersections and streets, tracks its position and uses a goal compass, and the failures trace to interference between superposed intersection features disrupting localization. Affordance packing, grouping intersections that share legal moves, limits the damage. The methodological point matters beyond navigation: behavioral failure is not evidence of absent representation, and the paper proposes mechanistic indicators to tell the two apart.
Source
↳ Follow the thread