Referential Security: AI Evaluations Are Broken Because Continuously Updated Models Violate Stable Identifier Assumptions
arXiv·medium signal
Proposes referential security as a new evaluation paradigm: current AI evaluations fail because model designations remain static while weights, prompts, retrieval mechanisms, and serving infrastructure undergo unannounced modifications. Any finding or audit attached to a model name becomes meaningless when the underlying artifact silently changes. Relevant to anyone relying on benchmark scores for deployment decisions.