Skills
An image patch on a web page made computer-use agents run a malicious terminal command, then finish the user's original task
AgentHijack tests image-triggered command injection end-to-end across screenshot input, VLM generation, action parsing and environment execution, deploying trained patches on author-controlled GitHub Pages and a locally hosted CSDN clone against five open-source or publicly available GUI-agent and VLM backends. Across 600 instance-level online cases, target-attack success reached 84.5% at the VLM output stage, 47.0% at action parsing, and 20.3% end-to-end with verifiable environmental consequences. Trajectory analysis found cases where the agent executed the injected terminal command and then carried on with the benign task, which is exactly the shape a user watching the screen would not catch.
↳ Follow the thread