Fetching from the wire…
Research2026-09-01 · source-backed
Measuring 8,234 revisions over nine years, with suppression detected semantically and validated against blinded hand labelling at 0.828 precision and 0.911 recall, exclusions were added 1,642 times and withdrawn 304 (arXiv 2608.31062). Per individual rule the ratio climbs to 13-to-1, and 86.7% are still in force three years later regardless of whether the rule is sole coverage for its ATT&CK technique. 31% of the narrowing is invisible to structural diff, and 64.1% of path-valued exclusions can be satisfied by an unprivileged process choosing a filename. Your detection coverage is quietly smaller than your rule count suggests.
Each link below shares sources, entities, or timing with this story.
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Measuring expert time on two datacenter GPU generations shows it is linear in neither token count (EPLB, LPLB, UltraEP) nor activated-expert count (METRO). Below roughly 156 to 168 tokens, HBM weight streaming dominates so cost attaches to activated replicas. Above it, grouped...
v3.5.12 carries 783 rules, up from 768. The release note describes the pipeline plainly: "tc-pr-back → safety gate → auto-merge → this release." GitHub The prior tagged release recorded 10/10 OWASP Agentic Top 10 coverage and PINT results of 62.7% recall at 99.7% precision, pl...
arXiv 2607.26791 gives agents a forensic disk snapshot of a breached host plus alerts, vuln scans, and baseline checks, then asks for intrusion, baseline-risk, and vulnerability-risk reports with a remediation plan. Ten cyber ranges, four entry-point types, 21 ATT&CK technique...
MLReproMutate applies controlled mutations across four classes (random seed, dependency pin, data split, cross-validation fold count) and runs them against the validation workflows the repositories already ship. Of 39 frozen repository-operator cases, 24 were evaluable with 23...
Single-shot prompting produced not one valid coverage-producing verification environment on the paper's benchmarks. AgentDV closes the loop with runnability filtering, CSR-grounded checking to cut hallucinated signals, and coverage-guided iteration against measured gaps. Using...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.