Fetching from the wire…
Research2026-09-07 · source-backed
LexFlip attacks the standard validity check for meaning-preservation metrics, which requires an identical pair to score highest and an unrelated pair lowest, and therefore passes any monotone function of token overlap (arXiv 2609.05296). The dataset releases 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of tokens. Seven embedding and BERTScore metrics spend 0.022 to 0.039 of their identical-to-unrelated range on such an edit, against 0.670 for bidirectional NLI, the one family the conventional check disqualifies. On FrJudge, a bare length feature outscores every semantic metric.
Each link below shares sources, entities, or timing with this story.
MAFIA (arXiv 2608.03844) targets the two conditions that describe production and that prior attacks failed against: large benign memory pools and active input auditing. It adds placement strategy (probe memory, allocate injection budget, schedule writes to stay retrieval-compe...
Multi-model systems treat different models as independent components even when failures stay correlated, and existing diversity metrics only capture differences in output meaning. The authors measure generative-process diversity via Normalised Compression Distance between raw...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Agents share transport and can call each other's tools but have no way to reconcile a fact phrased two ways (arXiv 2608.16357). Every incoming claim passes a five-outcome procedure (insert, merge, relate, conflict, reject) decided from scoped claim-key identity, embedding simi...
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
The attack iteratively pulls hidden chain-of-thought from black-box reasoning models using API-returned fidelity signals, reaching 66.4% near-verbatim extraction on open-source LRMs (trace length within 10% of target, 90%+ tokens matching exactly), generalizing to unseen datas...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.