SourcesDREAM Deep Research Evaluation with Agentic MetricsarXiv·medium signalXBlueskyLinkedInCopy linkAWS evaluation framework for deep research agents measuring multi-turn reasoning, tool use patterns, and research quality.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerHolding Back Ready Agent Turns Instead of Releasing Them Eagerly Cuts P95 Workflow Latency up to 3.50xarXiv 2609.10964Stack layer / ContrastSkill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmarkarXiv 2609.11682Stack layer / Follow-up threadAgent Benchmarks Carry a Double Measurement Confound: Moving Execution Decisions Off the Scaffold Turns a Flat Leaderboard Into a SpectrumarXiv 2609.09218Stack layer / Follow-up threadBenchmark Radar Ships a Daily-Updated Catalog of 1,283 AI Benchmarks With 12,916 Numeric Score ObservationsarXiv 2609.11115Stack layer / Follow-up threadScore retrieval sufficiency structurally before generation, not from the model's own confidencearXiv 2609.11023Stack layer / Follow-up threadMaP-WAM stores robot memory as completed segment records and turns them into plans, rather than replaying full historyarXiv (2609.11561)Stack layer / Threat patternClaude Code 2.1.269 Ships a Plugin Eval Runner and a Knob to Raise the Workflow Tool's Concurrent Agent Cap to 256Anthropic (claude-code CHANGELOG)Follow-up threadA survey and public compendium tries to fix the fact that 'agent' has no standard definition for evaluation purposesarXiv