Policy dependency / Contrast
A Fine-Tuned 4B Qwen in 2.6 GB Beats GPT-5.6 on a Transit-Kiosk Agent Benchmark, and PEFT Gains Vanish by 27B
arXiv 2609.10016
Stack layer / Threat pattern
AgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completion
arXiv
Stack layer / Threat pattern
Comments help LLM code generation only when they leak correct solution content, and comments from a different problem cut pass@1 by 20.8%
arXiv 2609.09242
Stack layer / Contrast
Frontier models now learn arbitrary ciphers from prompting alone, and encrypted harmful content slips past commercial classifiers as gibberish
arXiv 2609.09553
Stack layer / Contrast
A spec-first agent framework taxonomy: persuasion, front-loaded structure, or controls the agent cannot edit
arXiv 2609.09671
Stack layer / Follow-up thread
Φ-Bench tests whether LLMs can engineer their own serving and training stack, from kernels to end-to-end optimization
arXiv
Stack layer / Contrast
Ecdysis: fix the harness only for failure patterns that recur across tasks, not for each single failure
arXiv 2609.11677
Stack layer / Update thread
SynthID Watermarking in Claude Costs Three Points of Code Correctness on One Model, but Detection Is Near Chance
arXiv 2609.09604