Sources
AgentDrift: Tool-Augmented LLM Agents Show 65–93% Safety Failures Under Contaminated Tool Outputs — Standard NDCG Metrics Create Evaluation Blindness
Across seven LLMs (7B to frontier-scale) tested on 1,563 contaminated financial dialogue turns, risk-inappropriate recommendations appeared in 65–93% of turns while utility preservation remained at ~1.0 — meaning standard ranking metrics show no degradation. Narrative-only corruption (biased headlines, no numerical manipulation) triggers significant safety drift while evading consistency monitors. Safety-penalized NDCG drops utility to 0.51–0.74, exposing a critical measurement gap; none of the agents questioned tool reliability across 23-step conversations.
Source
↳ Follow the thread