Fetching from the wire…
Public story · 2026-08-10 · high
GPT-5.4-mini's attack success rate jumps from 41.7% to 72.9% when the same goal is split across three pages, per Borealis AI's StepJack benchmark.
Why now: The paper posted to arXiv in August 2026, with its dataset and code released for anyone testing an agent against page-hopping attacks.
Splitting an instruction across three pages hijacks computer-use agents more often, per Borealis AI's StepJack benchmark. Average attack success across models climbs from 31.3% at one page to 36.9% across three, and GPT-5.4-mini alone jumps from 41.7% to 72.9%. For anyone shipping an agent that clicks or browses on a user's behalf, that's the gap between a nuisance and a drained wallet or inbox.
The perverse part: EvoCUA-32B held up best in testing, not because it's more secure, but because it struggles to follow multi-hop reference chains. It can't reliably connect an instruction spread across pages, so it misses the attack along with the task it was supposed to do. Every other model tested got more exploitable as it got better at stitching multi-step context together.
StepJack runs 480 examples that break a malicious goal into steps placed along an agent's navigation path, instead of dropping it in one obvious block. Borealis AI has published the dataset and code alongside the paper.
The models best at multi-step reasoning are the same ones easiest to hijack this way, and that won't fix itself as reasoning gets better. It gets fixed, if it does, by an agent checking where an instruction actually came from before it acts. StepJack's numbers suggest nobody's shipping that check yet.
Each link below shares sources, entities, or timing with this story.
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Copilot uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; overlapping topics (agent, code).
GPT competes with Claude / Shared entity: Better / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover Better; reported by the same outlet (arxiv.org).
Copilot uses GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT; overlapping topics (agent, code).
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).