32 GPT-2 Runs Measure a Single Training Example: Learned in One Exposure, Undetectable by the Final Step
arXiv 2608.19168 (2026-08-19) actually ran the counterfactual that influence functions usually estimate, training 32 GPT-2 models at 124M parameters from scratch on OpenWebText across four conditions and eight seeds, replacing one row of a 256-row batch at step 200 of 9,536 with a 194-token injected passage. Fifty steps later the injected arm predicts the passage better by 0.039 and 0.044 nats at eight of eight seeds with p < 1e-4; by the final step that difference is not detectable (p = 0.25 and 0.71) against minimum detectable effects of 0.025 and 0.079 nats. Weight displacement reaches 44.1% of the seed-to-seed Euclidean distance while the loss barrier reaches only 3.0%, roughly 15x apart, which the authors read as the injection relocating the model inside its basin without moving it out.
↳ Follow the thread