Stack layer / Contrast
Recording an Agent Run at Its Non-Deterministic Boundaries Turns an Incident Into a CI Regression Test
arXiv 2609.20625
Stack layer / Threat pattern
A Model Can Fingerprint vLLM or SGLang From Its Own Output Tokens, Then Exploit It
arXiv 2609.20614
Stack layer / Threat pattern
Deobfuscating npm Packages Before the LLM Sees Them Raises Malicious-Detection Coverage From 69% to 99%
arXiv 2609.18862
Stack layer / Contrast
Adversarial Agents Ran Arbitrary Bash Past Claude Code Auto Mode and Codex Guardian in 79% of Trials
arXiv 2609.19587
Stack layer / Contrast
PACT benchmarks whether enterprise assistants break their own system-prompt rules under pressure from a persistent user
arXiv / HuggingFace Daily Papers
Stack layer / Contrast
The Most Collusive Pricing Model Honestly Reports Cooperative Intent, So Chain-of-Thought Monitoring Cannot Catch It
arXiv 2609.18346
Threat pattern / Contrast
All 15 Academic LLM Trading Agent Schemes Had Security Vulnerabilities and 80% Failed a Core Robustness Metric
arXiv 2609.19705
Stack layer / Update thread
Raising a Stated Failure Probability From 10% to 70% Changes Whether Frontier Models Check the Evidence by at Most 21 Points
arXiv 2609.17865