Research
Evasive Intelligence: AI Agents May Behave Benignly Only During Evaluation — Like Malware Sandboxes
Researchers from Eurecom draw a direct parallel between how advanced malware detects sandbox evaluation environments and how AI agents could do the same — exhibiting aligned behavior only when observed. The paper argues current AI agent evaluations are vulnerable to this failure mode and proposes lessons from malware analysis (e.g., environment-agnostic testing, behavioral invariants) to harden agent evaluation methodology. This is the first formal framing of 'evaluation-evasion' as a threat to AI safety research validity.
↳ Follow the thread