Rouxii: one prompt change raises an AI pentester's honeypot detection from 19% to 97%
arXiv·medium signal
arXiv 2609.26555 builds Rouxii, an autonomous pentesting framework that adds counter-deception to reconnaissance. Across three reasoning models, eleven network setups and 1,544 attack reports, matched cohorts that differ only in the prompt raise correct honeypot identification from 19% to 97% (OT services: 11% to 97%), with 0.7% false alarms. PentestGPT and HackingBuddy fail the same way, so defenders relying on honeypots to derail LLM attackers should assume that defense stops working once attackers add the prompt.