LLM Refinement of Decompiled Code Recovers Names From the Model's Prior, Not From the Decompiler Output
A Memorization Floor for LLM Refinement of Decompiled Code (arXiv 2609.17236, submitted 15 Sep 2026) uses a within-item control costing twenty API calls: refine a function, then refine it again from an input whose identifiers have been destroyed, and measure what survives. Recovery is real, refined output sits +0.072 to +0.137 above an arm-matched permutation null, but destroying the input's dataflow changes the naming gain by +0.001 (95% CI [-0.026, +0.026]). A second refiner from another vendor, registered in advance with byte-identical inputs, reproduced the result across twelve contrasts, and readability stayed at ceiling throughout so a human reader gets no signal that the names came from the prior.
↳ Follow the thread