Fetching from the wire…
OSS2026-09-18 · source-backed
arXiv 2609.20301 argues existing observability tools do per-execution debugging but not cross-run profiling, so nobody can answer where failures cluster or which tasks eat the budget at scale. The obstacle is that the responsible entity is a task intent like "diagnose authentication" rather than a code path with a stable identifier, so the authors define a semantic operation stack, recursively split trajectories at task boundaries, and emit pprof-compatible profiles. Segmentation reaches 0.764 B-cubed F1 against human annotations, and the profiles raise problem-localization MAP by up to 56% on three benchmarks.
Each link below shares sources, entities, or timing with this story.
First standardized benchmark for AI agent skills. 86 tasks, 11 domains, 7,308 test trajectories. Critical finding: curated skills +16.2%, self-generated skills +0%. Run it against your own skills. GitHub | Paper
It represents agents, sources, memories, claims and actions in a typed execution graph, traces ancestry to exclude permission-ineligible records, reranks by semantic similarity times path trust, and applies a risk-sensitive gate before execution (arXiv 2608.10509). Across 2,70...
RideWay pairs a stateful tool-calling ridehailing benchmark with Efficiency Utility, a success-gated metric discounting trajectories for excess tool calls and user-facing turns against task-specific reference effort, penalties calibrated from human paired preferences. Across 5...
Skill Issue points out that auto-synthesized skill documents get optimized against tasks a capable agent already solves with no document at all, leaving the optimizer nothing to measure. The fix is mining harder tasks by reverting merged PRs at a frozen base commit. On three K...
The benchmark hands a developer agent a real client engagement setup, business records, a requirements-holding client, a production API, an inherited codebase, cost and model limits, then scores it by deploying the customer-service agent it built against held-out simulated use...
arXiv 2608.02764 targets agents that issue refunds, reserve inventory and move money, where budgets and approval status change between authorization and effect. The authors define policy-state serializability: committed effects must be explainable as authorized against the pol...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.