Fetching from the wire…
Public story · 2026-07-22 · high
Asynchronous Attribution Fingerprint Vectors score campaigns at the proxy layer from tool-use timing and prompt residue, holding up under deliberate evasion tests.
Why now: The 2026-07-22 briefing flagged this for the evasion-resistant test results, not because multi-fleet attacks are suddenly common.
Attribution scoring pins coordinated attacks across separate AI agents at 0.82 pairwise AUC, per arXiv 2607.18826. Guardrails that evaluate one agent session at a time don't catch any of it. Real adversaries split a single campaign across independent agents and runtimes, so each local defense only ever sees a fragment.
The method scores campaign similarity with what the paper calls Asynchronous Attribution Fingerprint Vectors: tool-use patterns, timing, and prompt residue left across sessions. Against near-chance baselines, it hits 0.82 AUC. It held under controlled evasion, too.
Structural and stylometric residue in the prompts carried the strongest signal, ahead of timing alone, per the paper's ranking. The correlation itself runs at the proxy layer, outside any single agent's guardrail. No need to instrument every agent in a fleet one by one.
That's the bet worth making: per-agent guardrails turn into a wasted line item once you're running more than one fleet. The thing to watch: whether orchestration platforms ship this correlation themselves, instead of leaving each agent vendor to bolt on its own detection.
Each link below shares sources, entities, or timing with this story.
Agents share transport and can call each other's tools but have no way to reconcile a fact phrased two ways (arXiv 2608.16357). Every incoming claim passes a five-outcome procedure (insert, merge, relate, conflict, reject) decided from scoped claim-key identity, embedding simi...
Using CoderForge-Preview, described as the largest open dataset of coding agent trajectories, ensemble methods with SHAP attribution predict agent success before any run. Dominant drivers are patch fragmentation (how many places a fix has to touch) and repository scale. Prompt...
SkillsMetric evaluated 2,266 skills across 16 attack types, hitting F1 of 73.4%±0.5% overall (arXiv 2608.08468). Host destruction via shell commands: 0% detection. Natural-language prompt injection: 42%. If you lint third-party skills before install, this tells you precisely w...
Natural-language autoencoders judge an explanation of a hidden activation by whether the activation can be regenerated from it, a test structurally blind to individual false claims. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while only ~2%...
arXiv 2608.10314 had two LLM snapshots translate five theoretical accounts into code under structured-contract versus prose formats, producing 320 programs. Both primary hypotheses returned NOT_SUPPORTED, only 19 of 108 criterion evaluations passed, and cross-model format iden...
First large empirical study of static prompt-configuration files, across 11,427 repos, plus qualitative coding of 65 sampled files into a 65-code codebook (arXiv 2608.10622). Adoption emerged fast from mid-2024 but clusters in small, low-activity, single-maintainer repos. Cont...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.