A local 8B model closed the full observe-decide-act loop of a remote access trojan — and completed only 10.9% of the checklist
arXiv 2608.03009 (Aug 4) builds an agentic RAT in a network-isolated lab from Kali Linux, Metasploitable2, LM Studio and a local 8-billion-parameter Dolphin-family model, then asks whether a model that small can drive autonomous intrusion with no cloud service and no operator in the loop. It can, architecturally: the SLM interpreted ranked reconnaissance evidence, selected actions and obtained verified root-shell access on real vulnerable services. It is not yet operationally reliable — the model hallucinated commands, misread output, recovered inconsistently, and cleared just 10.9% of a deliberately strict checklist. The authors frame that gap as a property of today's small models rather than a ceiling, which is the part defenders should plan around.
Source
↳ Follow the thread