Fetching from the wire…
Public story · 2026-09-06 · high
The paper doesn't say which published results the practice undermines, only that the comparisons can't be trusted.
Why now: The critique posted to arXiv in September 2026.
Side-channel attack researchers compare new exploits against numbers borrowed from earlier papers, rather than retesting them, per a paper on side-channel benchmarking.
These attacks hinge on exact hardware and cache setup, and a small difference between machines can swing results enough to make comparisons meaningless. The proxies show up at top-tier security venues, the ones setting the bar for what counts as progress.
Treating a borrowed number as an apples-to-apples comparison risks crediting a new attack with an improvement that's really just a different test machine.
The proxies include covert-channel bandwidth and key recovery against naive AES and RSA implementations, standing in for a retest on the attack's own hardware.
The paper is a methodology critique, not a new attack, which is rarer in a field that tends to reward novel exploits over process work. It doesn't name which published results the comparisons undermine or propose a replacement benchmark, so the diagnosis arrives without a fix.
Each link below shares sources, entities, or timing with this story.
The FSE '26 paper argues SWE-bench, SWT-bench, and AgentBench capture narrow synthetic slices, and proposes contamination-aware, trajectory-aware, in-the-wild evaluation using agents' commit signatures to study real vs human contributions over time. (arXiv) Pair this with the...
RealSWE builds a six-category information taxonomy and four style dimensions, then compares real prompts from SWE-chat against SWE-bench Verified and Pro. Also: 87% of real prompts are casually written, against 94% of benchmark problems written formally. They release 381 multi...
Multiple independent models train against each other with peer-derived rewards and no ground-truth labels, gaining 3.0-8.6% across seven text benchmarks and 2.3-7.2% across four multimodal ones. The mechanism claim matters more than the numbers: varying architectures, model si...
The standard protocol ablates a latent and measures effect at the token where it fires hardest, but that token is chosen by the dictionary under evaluation. Two dictionaries get compared at different places. Training six autoencoders from one initialization showed 7.6% and 11....
arXiv 2608.13010 scores top-five retrieval candidates against ranks 6–20 of the same query to spot answer-anchor concentration, and separately compares documents to lexically distinct neighbors to catch coordinated density before any query arrives. Deployed jointly, attack suc...
MAFIA (arXiv 2608.03844) targets the two conditions that describe production and that prior attacks failed against: large benign memory pools and active input auditing. It adds placement strategy (probe memory, allocate injection budget, schedule writes to stay retrieval-compe...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.