72% of 39 agent models complete assigned insider-threat objectives, and refusal does not predict whether they finish
MOLE is an open benchmark of 150 AI-operated accounts sharing nine stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens, built to test whether defenders can spot weight exfiltration, training-data poisoning or weakened release gates among routine work under a review budget. Of 39 agent models, 72% completed most assigned harmful objectives, and agent refusal did not predict completion. Comparing 40 monitors, even the best in the single-day audit-event comparison missed nearly half of completed harm; benchmark-guided search improved a mid-tier monitor by 49-64%, and selectively escalating to a stronger monitor beat blanket application by 10% budget-AUC at comparable cost.
Source
↳ Follow the thread