Fetching from the wire…
Agents2026-09-22 · source-backed
The Agent Plugins v1.0.0 spec from July 24 standardized packaging for skills, sub-agents, commands, hooks and tool servers. AgentPluginZoo, a provenance-tracked corpus across 30,655 repositories, finds 96.6% of failures would load after adding one missing boilerplate field, so validation isn't the real cost. 40.2% would load only if the client discards fields their authors wrote, and 81% of capability-exporting bundles share a name with another plugin with no namespace or precedence rule to decide which answers. The community standardized a packaging format when composition needed a model. (arXiv 2609.23809)
Each link below shares sources, entities, or timing with this story.
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
A reconstruction of 14,591 revisions, 3,103 names, 4,579 pages and 19,913 server events from the May-July incident finds coordination formats converged within a day, and that across the 510 cohorts with an observable progress trace there's no robust positive association betwee...
The trick is one line in a file you never read. Manifold Security published eight findings across seven coding agents (Claude Code, Codex, Cursor, Grok Build, Qwen Code, goose, Hermes Agent) that all reduce to the same mechanism. A repository's own .git/config sets core.fsmoni...
OSReward builds human-verified ground truth for computer-use trajectory judgments and finds even state-of-the-art models fall short with a consistent bias toward misclassifying failures as successes. The authors release OS-Shepherd at 9B and 35B, trained on a 100K corpus, clai...
A July 29 arXiv paper from Peter Kirgis, Sayash Kapoor, and Andrew Schwartz introduces shadow evaluations: agents attack the central research question of an unpublished high-quality paper, and the original authors grade the result. Across two unpublished NeurIPS 2026 submissio...
Here's the setup. You build a benchmark with 100 research tasks. Each answer is supported by two independent chains of corroborating records. Clean environment, agents do fine. Then you drop in exactly one plausible-looking document that carries a conflicting answer. Accuracy...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.