Fetching from the wire…
Security2026-09-05 · source-backed
arXiv 2609.04058 reports deployed silicon where the standard acceptance gate structurally cannot detect a class of ML-DSA defects: signing resamples until a candidate meets its norm bounds, so the executed path varies with the message, while known-answer tests use fixed vectors and only reach the depths their seeds trigger. The accelerator carried a norm check that outran block-RAM latency, leaving each candidate's final coefficients unverified. Replacing the gate with a byte-exact golden-reference oracle plus randomized adversarial soak gave 301,343 data-dependent signings with zero escapes. This is the "you have not watched your test fail" problem in hardware (arXiv).
Each link below shares sources, entities, or timing with this story.
Ockhamareto (arXiv 2608.24473) reinforces a unit-test rollout only when it's non-dominated on both mutation-killing and test count, then ties each test's killing power back to specific source tokens. Against MIST-RL that's a 3.4x better per-test trade-off, plus 30 to 35 percen...
Destefanis and Aste modeled 1,902 multi-agent AI coding runs as temporal networks of agents, files, and timestamped messages (arXiv 2608.16801). This is the most useful paper in today's set and it lands directly on top of what everyone shipped this week. Three results. Direct...
When an agent consolidates an external observation into long-term memory, attach platform-controlled metadata recording the source's trust level, then gate tool execution by matching action risk against supporting-memory authority. Laundered memories hit a 1.000 attack success...
SARC-DQ found competent agents converted freshness/lineage/provenance defects into costly actions about 60% of the time, with both data-quality flags and the agents' own hedging detecting them at chance. The conversion rate was flat across four model tiers spanning a 15x price...
Data that contradicts the vibe. That's rare enough to lead with. Dipongkor, Baral, Lam and Moran analyzed 4,882 pull requests from five coding agents in the AIDev dataset (532 Java, 4,350 Python), accepted to ICSME 2026. The findings, in order of how much they should change yo...
It's a benchmark of 56 contract-defined backend tasks, judged only through black-box HTTP tests against an OpenAPI contract, so there's nowhere to hide (arXiv). GPT-5.5, the best model, succeeds on 55.4% under the base oracle and drops to 28.6% under the final hardened oracle....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.