Fetching from the wire…
Research2026-09-25 · source-backed
arXiv 2609.30120 re-examined 64 reports on 16 migration tasks for a shipped coding-agent skill. Raw gain was 93.83 to 98.75, concentrated in a single task. Hand review found grading errors in both directions, including a containment check awarding full credit for accepting the parent directory. Corrected decisions moved the confidence interval to or across zero. Judges from two other model families agreed on 91.8% and 95.7% of decisions and still estimated gains of 10.63 and 6.09 points. Your skill's measured benefit depends substantially on which model graded it.
Each link below shares sources, entities, or timing with this story.
This one should change how you read leaderboards. A physics benchmark audit put faculty and graduate researchers through six widely used physics benchmarks, including ones feeding the Artificial Analysis Intelligence Index that half the industry quotes. They reviewed problem s...
ProgramBench dropped a benchmark that should make every "AI will replace developers" hot take age badly. The setup: give an agent a compiled executable and documentation, then ask it to architect and implement a complete codebase that reproduces the original program's behavior...
Recuris (arXiv 2608.24876) keeps a Working Memory tracking current task progress separate from an Experiential Memory of learned skills, so skill selection indexes against what the task needs now rather than the whole history. It improves 35 of 37 model-benchmark pairs, gains...
arXiv 2609.26550 (CMU, submitted September 22) compares Jev against sixteen generative and reward-model judges under blinded human adjudication. It costs $0.044 per 1,000 judgments at 152ms median and lands within 3 points of the strongest LLM judge on RewardBench-style prefer...
Across 30 models from three families, verbalized confidence compared against logits-based confidence on 8 classification tasks and semantic entropy on 2 generation tasks: instance-level association is weak on average, improving only on easier items and stronger base models. In...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.