Fetching from the wire…
Public story · 2026-09-18 · high
A resampling method cut eval costs 58% and scored higher using only the fraction of tasks that predict the full benchmark.
Why now: Eval budgets climb every time a benchmark suite grows, and DeltaSelect's DeepSWE numbers show most of that spend buys no signal.
DeltaSelect resamples DeepSWE's published trials and finds that only 22 of its 113 tasks track full-benchmark results closely enough to trust. The paper's bar is a fifth-percentile Pearson correlation of at least 0.50, a threshold just 19.5% of tasks clear. The other 91 tasks don't reliably tell you whether a candidate change to a model or prompt helped.
DeltaSelect also works as a method. It picks the subset of tasks whose single-run results predict the full benchmark, then fits that fixed set to a dollar budget. Applied in a case study on gpt-5.6-luna, it drove skill and instruction revisions across 13 evaluations for $27.86 total.
The eval that came out the other side ran 58.1% cheaper per run, $1.75 against $4.18 (p=0.008), and scored higher on the calibrated metric, 42.36% against 36.46%. Teams running the full 113-task suite on every candidate change are paying for 80% of that spend on tasks the paper says carry no signal. The fix is picking the fifth of the suite that predicts the whole benchmark, not adding more tasks to it.
Each link below shares sources, entities, or timing with this story.
This is the week's most useful number, and it took an actual experiment to produce it rather than a launch post. Together AI ran 113 DeepSWE tasks at 4 trials per config: 452 GLM-5.3 rollouts and 448 GLM-5.3-Flash rollouts. GLM-5.3 scored 69.0% pass@1 at $3.99 per rollout. Fla...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
Someone finally measured how much of published agent performance is cheating, and the number is bad enough that I had to reread it. Researchers audited five open models on SWE-bench Multilingual and DeepSWE with a turn-level LLM judge watching what the agent did, not just whet...
Simon Willison found it in the OpenRouter price list. 1.6T MoE, ~49B active, 1M context, up to 384K output, at $0.435/M input (cache miss) and $0.87/M output. The agentic-coding deltas versus the preview are the story: DeepSWE 12.8 → 62.7, CyberGym 52.7 → 83.3, Terminal Bench...
The changelog lists Terminal Bench 2.1 at 82.7, NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon verified 70.3, DSBench-FullStack 68.7 and DSBench-Hard 59.6: figures DeepSeek says far exceed V4-Pro-Preview. Native Responses API support and specific Codex adaptation. Only...
arXiv 2607.27146 attacks from-scratch program synthesis, where agents get only natural-language docs and an execute-only binary as oracle. The pipeline auto-converts open-source command-line programs into source-free training environments and uses GLM-5.2 as teacher for synthe...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.