Fetching from the wire…
Public story · 2026-09-26 · high
Confidence intervals came out narrower in 15 of 16 model panels once every task got covered before any got a second run.
Why now: The bound and its LiveCodeBench replay went up September 26.
Spreading a fixed eval budget across every task before resampling any of them cuts score error by 87% compared to pooling extra runs on fewer tasks, according to a paper posted September 26 that derives bounds on how much replication a repeated eval actually needs.
That's a big swing in how trustworthy a leaderboard number is for the same total compute spend. Teams comparing models on a budget can get a tighter or looser confidence interval purely from how they allocated samples, before any model result changes.
The test: 16 models, 880 LiveCodeBench tasks, five outputs per task, identical total budget under both designs. One design touches every task at least once before spending anything on a second sample. The other pools samples uniformly, which in practice means some tasks get hit five times while others get skipped. Coverage-first cut median point-estimate error (MSE) by 87.0%. Interval width came down too, median 30.6% narrower, and coverage won in 15 of the 16 model panels.
The paper doesn't say whether the bound holds outside this setup. LiveCodeBench is one benchmark family scored by pass@k, and 880 tasks is one scale. It doesn't test benchmarks with uneven task difficulty or scoring that isn't pass@k, and whether the 87% figure holds at 88 tasks or 8,800 is untested here.
If you're running repeated evals on a fixed budget, the fix costs nothing extra: touch every task once before any task gets a second sample.
Each link below shares sources, entities, or timing with this story.
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
Across 614 problems from APPS, HumanEval+ and LiveCodeBench, hierarchical collaboration was worth 2.4 pass@1 points on the easiest problems and 21.1 on the hardest, at a flat ~10x token cost throughout (arXiv 2609.13890). DATS predicts each topology's success probability and p...
Across 2,700 runs with Kimi K3 at three thinking-effort levels, stripping a full specification down to a user story increased token spend by 29.7%, while no prompt change affected run-to-run variance at all (arXiv 2608.25399). Sensitivity is strongly task-dependent, ranging 13...
A 15-run pilot, a pre-registered 20-run confirmatory ablation and a pre-registered 2x2 factorial with 40 runs across two vulnerable lab systems (arXiv 2609.15887). Removing verification raised reported findings (median 2 against 0, p = 0.00003) and cut precision (0.353 against...
Announced August 31, exposing SFT, DPO, reward finetuning and RL across three surfaces (Fireworks). Harvey post-trained Kimi K3 into "Harvey Tenet" scoring 19.7% all-pass on LAB against 10.8% for the base model at comparable cost. Vercel reports a 93% error-free generation rat...
Wafer.ai ran K3 at TP8 on a single MI355X node: 952 tok/s aggregate, 118 tok/s single-stream decode, ~13k tok/s steady-state prefill after fixing missing PyTorch sampling functions and AITER MLA prefill kernels via SGLang and ROCm. Against a two-node TP16 B200 deployment that'...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.