Fetching from the wire…
Agents2026-09-25 · source-backed
arXiv 2609.30137 runs the Snowglobe simulator against Nubank's Card Delivery and Card Management chat agents in Brazil. Simulated evaluator scores tracked production across four deployed versions. Simulation-guided iteration raised transactional NPS by 36.69 points in a live A/B test, and an open-weight configuration chosen from over 16,000 simulated conversations lifted self-service rate 8.82 points to Nubank's highest recorded level with no significant tNPS change. This is one of very few production-scale reports where pre-deployment simulation demonstrably predicted live results in a regulated industry. If you have a high-volume conversational agent and no simulator, this is the paper to bring to whoever controls the budget.
Each link below shares sources, entities, or timing with this story.
arXiv 2609.20301 argues existing observability tools do per-execution debugging but not cross-run profiling, so nobody can answer where failures cluster or which tasks eat the budget at scale. The obstacle is that the responsible entity is a task intent like "diagnose authenti...
arXiv 2609.10263 separates what a persistent agent stores from what it uses, because a superseded fact misleads a current-state answer while remaining necessary for a historical query. A retained archive holds everything; a query-conditioned view governs influence, with same-s...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
arXiv 2608.02764 targets agents that issue refunds, reserve inventory and move money, where budgets and approval status change between authorization and effect. The authors define policy-state serializability: committed effects must be explainable as authorized against the pol...
ScrambleToolBench strips semantic cues from tool schemas, then injects mapping drift, stochastic failures, and temporal execution windows. Frontier models discover the initial mapping fine but show belief inertia or fall back to exhaustive search under structural change, and i...
David Vélez and Robin Vince joined both the OpenAI Foundation and OpenAI Group PBC (OpenAI). Both skew heavily toward finance and regulated-industry governance rather than AI research or product. Read alongside this month's IPO reporting, the board composition looks like publi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.