Surrogate Fidelity asks when open LLMs can be used to explain closed ones
arXiv·low signal
Mechanistic interpretability needs full model internals, but most deployed models expose only output log-probabilities, creating a surrogate problem: when do measurements on open models license claims about a closed model? The paper evaluates surrogate fidelity at prediction, attribution, and representation levels, finding log-odds give an API-compatible scalar readout for binary tasks. Useful framing for practitioners auditing closed APIs they can't inspect directly.