Attackers don't need your credentials - they can just spend your agent's resources: 84.06% success, best defense still leaves 55.11%
Resource hijacking is a distinct attack class from data exfiltration: the attacker induces an agent to invoke, consume, transfer, or control high-value resources for the attacker's benefit without ever obtaining the resource or its credentials - think burning your compute budget, triggering paid API calls, or driving organizational workflows. ResourceHijackBench supplies 300 attack scenarios across six resource categories in isolated local environments that record actual resource consumption rather than simulated outcomes. Undefended, OpenClaw showed an average 84.06% attack success rate, and the strongest evaluated defense still left 55.11% - so budget caps and per-tool spend ceilings enforced outside the model, not refusal training, are the layer that matters here.
↳ Follow the thread