Fetching from the wire…
Agents2026-09-17 · source-backed
RideWay pairs a stateful tool-calling ridehailing benchmark with Efficiency Utility, a success-gated metric discounting trajectories for excess tool calls and user-facing turns against task-specific reference effort, penalties calibrated from human paired preferences. Across 58 tasks and 24 models the fitted penalty for excess turns is about 2x that for excess tool calls. On held-out preferences the metric hits 78.7% accuracy overall, 90.6% when trajectories differ in turns, and chance level when they differ only in tool calls. Asking the user one extra question costs you more than two extra searches.
Each link below shares sources, entities, or timing with this story.
Across 30 models from three families, verbalized confidence compared against logits-based confidence on 8 classification tasks and semantic entropy on 2 generation tasks: instance-level association is weak on average, improving only on easier items and stronger base models. In...
Studdiford and Lupyan tested human participants and 25 LLMs on common-sense causal reasoning and found shared, predictable error patterns triggered by irrelevant prompt details (arXiv). They localized the attention heads driving it. The uncomfortable implication: the gap betwe...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
Three models were evaluated on 832 Defects4J bugs with hallucination tracked across final patches and the intermediate artifacts guiding them (arXiv 2609.04909). Only 21.0% to 55.9% of generated patches passed the developer-written test suite, and manual analysis of 812 sample...
ISM shapes the semantic relationship between a user prompt and a skill's metadata so the router picks the attacker's skill, with no steering text anywhere in the payload. Across four task domains and eight selector models it raised average target-selection rate from 15.2% to 6...
If you're on Pro, Max, or Team, the permission prompt you've been hitting Enter on for a year goes away Friday. Anthropic confirmed auto mode becomes the default, replacing per-call approval with a classifier that inspects each tool call for irreversible, destructive, or out-o...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.