Prompt Injection Hidden in Log Entries Makes SOC LLMs Call Compromised Traces Benign — but Their Own Explanations Leak the Attack
'Just Testing, Move Along' (arXiv 2607.24174, July 27) builds a framework for evaluating prompt-injection attacks against LLM-based system-log interpretation in Security Operations Center workflows, generating adversarial examples from real cyber-attack log traces via generic injection generation, refinement, and attack-specific optimization. Across multiple state-of-the-art LLMs, injected log entries caused malicious traces containing clear indicators of compromise to be classified as benign. The useful counterpart finding: the natural-language explanations the models emit alongside their classifications frequently contain traces of the adversarial manipulation, giving defenders a cheap detection channel that requires no model changes — just don't discard the rationale.
↳ Follow the thread