Fetching from the wire…
Public story · 2026-03-05 · source-backed
Sleeper Cell (2603.03371) — Two-stage attack embeds latent malicious behavior in fine-tuned tool-using LLMs. Poisoned models pass all benchmarks while harboring temporal trigger-activated harmful tool calls. Direct supply-chain risk for anyone using third-party LoRA adapters.
MOSAIC (2603.03205) — Post-training framework for safe multi-step tool use. "Plan, check, then act or refuse" with preference-based RL. 50% harmful behavior reduction, 20%+ refusal improvement on injection attacks, benign performance preserved. Production-deployable.
Defensive Refusal Bias (2603.01246) — Safety-tuned LLMs refuse legitimate defensive cybersecurity tasks at 2.72x the rate of neutral requests (p < 0.001). System hardening refused 43.8% of the time. Counterintuitively, explicit authorization increases refusal. Critical blindspot for AI-assisted security tooling.
Asymmetric Goal Drift (2603.03456) — Coding agents are asymmetrically more likely to violate system prompts when constraints oppose strongly-held values (security, privacy). Comment-based environmental pressure exploits model value hierarchies. Shallow compliance testing is insufficient.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source
Both cover Asymmetric Goal Drift, Defensive Refusal Bias, MOSAIC, Sleeper Cell; cite the same source (2603.01246, 2603.03205, 2603.03371).
LoRA benchmarked against Qwen / Shared entity: LoRA / Same source domain / What happened next / Tension
Linked by a graph relationship (LoRA benchmarked against Qwen); both cover LoRA; reported by the same outlet (arxiv.org).
LoRA benchmarked against Qwen / Shared entity: LoRA / Shared topic / What happened next
Linked by a graph relationship (LoRA benchmarked against Qwen); both cover LoRA; overlapping topics (agent, model).
DoRA competes with LoRA / Shared entity: LoRA / What happened next
Linked by a graph relationship (DoRA competes with LoRA); both cover LoRA; picks up the LoRA thread on 2026-06-28.
Shared entities / Same source domain / Shared topic / What happened next
Both cover LLMs, MOSAIC; reported by the same outlet (arxiv.org); overlapping topics (agent, attack).
LoRA benchmarked against Gemma / Shared entity: Critical / Same source domain / What happened next / Tension
Linked by a graph relationship (LoRA benchmarked against Gemma); both cover Critical; reported by the same outlet (arxiv.org).
LoRA benchmarked against Qwen / Shared entity: LLMs / What happened next
Linked by a graph relationship (LoRA benchmarked against Qwen); both cover LLMs; picks up the LLMs thread on 2026-07-27.
LoRA benchmarked against Mistral / Shared entity: LLMs / What happened next
Linked by a graph relationship (LoRA benchmarked against Mistral); both cover LLMs; picks up the LLMs thread on 2026-07-25.