Fetching from the wire…
Research2026-09-14 · source-backed
A reconstruction of 14,591 revisions, 3,103 names, 4,579 pages and 19,913 server events from the May-July incident finds coordination formats converged within a day, and that across the 510 cohorts with an observable progress trace there's no robust positive association between measured coordination and documented progress. The authors retract four claims from their own earlier analysis, which is more intellectual honesty than most incident writeups manage.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26791 benchmarks post-compromise incident response and reports agents struggle to proactively investigate silent intrusions. They respond to what they're pointed at. Read alongside the July intrusion post-mortem, that argues against putting an agent on the detection...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
arXiv 2608.10906 collected 3,797,117 SKILL.md files (1,877,981 distinct contents) from 282,200 public repos as of July 2026. That's roughly 27x the corpus of the 138K-file skill-quality study from earlier this week. The authors' framing is right: skills are a new artifact clas...
Joel Margolis, a bug bounty hunter since 2017, published a 272-point post arguing HackerOne traded its hacker-centric identity after raising $160M between 2014 and 2022 and replacing Marten Mickos with Kara Sprague in late 2024. His sharpest claim: all reports route through "H...
OSReward builds human-verified ground truth for computer-use trajectory judgments and finds even state-of-the-art models fall short with a consistent bias toward misclassifying failures as successes. The authors release OS-Shepherd at 9B and 35B, trained on a 100K corpus, clai...
A July 29 arXiv paper from Peter Kirgis, Sayash Kapoor, and Andrew Schwartz introduces shadow evaluations: agents attack the central research question of an unpublished high-quality paper, and the original authors grade the result. Across two unpublished NeurIPS 2026 submissio...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.