Fetching from the wire…
Infra2026-09-17 · source-backed
arXiv 2609.18849 proves no estimate fixed before a tool call starts can even rank calls correctly, which kills the whole family of guessing from tool name, history, declared duration or engine occupancy. A census of four public agent corpora found a readable progress signal in most tool time once the tool is allowed to emit it, as fraction-of-work-remaining or an end-is-near flag. At the moments a KV-cache decision gets made, reported progress is several times to an order of magnitude more accurate than the best published predictors, holds under environment change, and cut p90 time-to-first-token in a production engine through a few small hints. If you own the tools, emitting progress is nearly free and it beats every predictor built on top of you.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
The failure mode is a well-formed but policy-forbidden call, cancel a booking, change a passenger count, that neither the tool nor the agent's self-report flags (arXiv). In the airline domain tested, the fix wasn't more reasoning. It was cheap, read-only deterministic gates th...
arXiv 2609.18674 extends CaMeL with a static verification layer. CaMeLoT translates a generated plan into a finite-state transition system labeled with tool calls, provenance and taint, then checks it against CTL policies with nuXmv before execution starts. Unsafe plans get re...
RideWay pairs a stateful tool-calling ridehailing benchmark with Efficiency Utility, a success-gated metric discounting trajectories for excess tool calls and user-facing turns against task-specific reference effort, penalties calibrated from human paired preferences. Across 5...
arXiv 2609.17698 read documentation, source, config and tests across 157 LLM-agent projects with 100+ stars. Coverage fragments in a specific, checkable way: a guard on the direct tool call and nothing on the shell that reaches the same effect, tests that rarely exercise bound...
A multi-tenant tool that accepts a tenant identifier and validates it against the caller's entitlement is still delegating resource selection to a process whose context an attacker may control (arXiv 2609.14780). In a 373-trial ablation across eight model configs and two trans...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.