Fetching from the wire…
Research2026-09-17 · source-backed
XConf breaks the premise every existing confidence estimator shares, that confidence is readable off the current inference through introspection, token probabilities or resampling. It stores graded past episodes holding the task, reflection, stated confidence, outcome and a lesson, recalls similar tasks met with similar stated confidence to extract real success rates, then has the model name its recurring failure pattern and revise. Nine benchmarks, four models, three families: matched or beat ten-sample self-consistency on AUROC in 23 of 24 comparisons, about a tenth the ECE and a tenth the generation cost, with selective prediction raising delivered agent success up to 8.7 points. Anyone routing work by confidence should read this one.
Each link below shares sources, entities, or timing with this story.
Across 2,823 committed episodes on three frameworks, a one-class echo-state-network ensemble with CUSUM alarms catches 71% of mid-episode failures at a 5% false-alarm budget, three orders of magnitude cheaper than a judge call. But learned monitors don't transfer (AUROC 0.527...
Agents get only a CWE description and terminal access, and have to find the implementing files across 500 real vulnerabilities from 290 repositories, six package ecosystems and 147 CWE categories (arXiv 2609.15939). Across 27 language models and four static-analysis tools on a...
Attnlocate (arXiv 2608.24022) aggregates attention across heads and layers into a token-level feature space, then runs a 1-D U-Net with an anchor-free detection head to find the traces behavior-guiding instructions leave behind, adjudicating the tool call based on the authorit...
Researchers loaded five systems with a revoked policy and its replacement, then measured retrieval and downstream action across nine policy scenarios, nine models and six defense conditions. Wherever the revocation label was visible to the retrieval layer, the revoked fact cam...
A June 26 paper (arXiv:2606.26294) describes a self-improving architecture where the agent and the evaluator that scores it evolve together, specifically to avoid the stagnation of optimizing against a fixed, gameable reward. (arXiv) Anyone building a self-improving harness ha...
arXiv 2608.24358 switched models mid-run on long coding tasks using cheap/expensive pairs from the Claude and GPT families. Full-trajectory escalation from weak to strong recovers under half the gap while costing a substantial premium, which the authors call the handoff tax. D...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.