Fetching from the wire…
Public story · 2026-08-19 · high
Even the best tested defense against resource-hijacking attacks still let more than half of them through, per a new benchmark.
Why now: arXiv posted the ResourceHijackBench paper, 2608.15108, with results specific enough to act on as of August 19.
ResourceHijackBench found AI agents get hijacked into spending resources for an attacker's benefit, without the attacker ever touching a credential, per arXiv paper 2608.15108.
Undefended, agents handed attackers what they wanted 84.06% of the time across 300 scenarios in six resource categories. Think compute budgets burned, paid API calls triggered, workflows steered toward someone else's outcome.
That's a different threat than data exfiltration. The attacker never steals your API key. They trick the agent into using it, and the agent keeps full authority to act, just now on the attacker's behalf.
The 300 scenarios ran in isolated local environments. Researchers measured actual resource consumption, not simulated outcomes.
The strongest defense the researchers tested still let attackers through 55.11% of the time. Refusal training and prompt-level guardrails aren't closing this gap on their own.
The paper points at a fix that sits outside the model: budget caps and per-tool spend ceilings enforced independently of what the model decides. A related report on composing agent guardrails as hard policy rather than trained-in refusal points at the same direction.
If you're running agents with access to anything that costs money or moves state, the question isn't whether the model refuses a bad request. It's whether something outside the model can say no regardless.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.80).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.77).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.77).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.76).
Same source
Cite the same source (arXiv 2608.15108 - Beyond Direct Access: Resource Hijacking in LLM Agents).
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).