Fetching from the wire…
Public story · 2026-08-03 · high
Preprocessing and labeling choices differ paper to paper on the same two datasets, so published rankings may not hold.
Why now: The comparison lands as security teams keep picking lateral-movement detectors off leaderboard tables built from these same two datasets.
Three well-known lateral-movement detectors scored differently once a paper reran them under matched preprocessing rules, per arXiv 2607.29390. That's a problem for anyone who picked a detection tool off a benchmark table built from these two datasets.
The paper targets methodology, not the detectors themselves. On the two datasets most lateral-movement research relies on, preprocessing and event-labeling methodology varies substantially paper to paper. That choice quietly decides how fair the resulting evaluation is. Standardize it and the same detectors don't land where their original papers put them.
The paper doesn't say which direction each detector moved, or name the three detectors specifically. That's a gap if you're trying to map this straight to a purchase decision, but it doesn't undercut the methodology finding.
The fix isn't a better detector, it's a standard labeling spec for these two datasets that every paper has to use. Watch whether that spec shows up. Until it does, every leaderboard built on these datasets stays provisional.
Each link below shares sources, entities, or timing with this story.
Sampled softmax cuts the O(nK) memory of full-vocabulary classification to O(nk), but for fixed budget B = n·k it's been unclear whether to buy batch or negatives (arXiv 2608.11061). Analyzing convergence under standard smoothness and variance assumptions, the fastest converge...
Learned KV eviction has a soft-to-hard mismatch: training uses differentiable gates that attenuate contributions, but inference only saves memory when entries are physically removed (arXiv 2608.23296). A controlled 2x2x2 over attention type, learned gating and positional encod...
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
arXiv 2608.03609 formalizes agentic systems over relational data as Stateful Tool-Enabled Agentic Deployments and proves verification against First-Order CTL specs is undecidable. Under a finite-domain restriction it becomes PSPACE-complete, but only if renaming opaque identif...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
The FSE '26 paper argues SWE-bench, SWT-bench, and AgentBench capture narrow synthetic slices, and proposes contamination-aware, trajectory-aware, in-the-wild evaluation using agents' commit signatures to study real vs human contributions over time. (arXiv) Pair this with the...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.