Fetching from the wire…
Agents2026-09-15 · source-backed
This paper builds an effect-history model separating events in the world from the runtime's observations of them, then catalogs eight recurring anomalies at the agent-tool boundary under retries, speculative execution, concurrency and partial failures (arXiv 2609.15397). Missing effects. Duplicated effects. Aborted effects that survive. Committed effects depending on provisional state later withdrawn. Measuring the standard annotation vocabulary across 98,291 tools, the fields get emitted widely but give only coarse call-level hints, and none of the required boundary capabilities is expressible. Four of the eight anomalies cannot be excluded by black-box tool invocation at all.
Each link below shares sources, entities, or timing with this story.
Measuring 8,234 revisions over nine years, with suppression detected semantically and validated against blinded hand labelling at 0.828 precision and 0.911 recall, exclusions were added 1,642 times and withdrawn 304 (arXiv 2608.31062). Per individual rule the ratio climbs to 1...
Measuring effective stream counts, cross-stream residual weights and inter-stream cosine similarity across the four-stream mHC residual pathway shows a typical attention or FFN site effectively uses about two streams, and layers 22-42 mostly carry each stream forward separatel...
arXiv 2609.04172 trains OPD with a single query and finds it keeps improving for hundreds of steps across task domains and model families. Measuring state coverage, the fraction of full-data states a query set's rollouts reach, one query hits 71.5% and 16 semantically distinct...
Measuring teacher supervision during on-policy distillation shows substantial noise that worsens as the teacher scales, yet the student converges comparably whether that supervision is kept or stripped (arXiv 2608.31046). Learning concentrates on low log-probability tokens, an...
Measuring expert time on two datacenter GPU generations shows it is linear in neither token count (EPLB, LPLB, UltraEP) nor activated-expert count (METRO). Below roughly 156 to 168 tokens, HBM weight streaming dominates so cost attaches to activated replicas. Above it, grouped...
Nolan Lawson (ex-Microsoft, ex-Salesforce) published an essay that hit 662 points and 247 comments on Hacker News. His argument: stop using LLMs to ship faster. Use them to ship better. His approach runs multiple models to review code, ranks findings by criticality, and filter...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.