Fetching from the wire…
Research2026-09-25 · source-backed
arXiv 2609.30074 bootstrapped its own ranking across 8B to 675B models with caching disabled. The two worst models held rank in 99% and 86% of replicates, the middle four in 27 to 48%, the top two in 68%. Two defensible rules for merging runs changed four of eight rows. Four of eight API endpoints were withdrawn inside ten weeks, so the study can't be rerun. Small-prompt-set leaderboards should be reporting rank stability, and almost none do.
Each link below shares sources, entities, or timing with this story.
535 participants solved a 40-item reasoning battery either unaided or while required to consult GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash or Kimi K3, with each model also answering every item alone 100 times under matched elicitation. In a reference comparison, roughly h...
This one should change how you read leaderboards. A physics benchmark audit put faculty and graduate researchers through six widely used physics benchmarks, including ones feeding the Artificial Analysis Intelligence Index that half the industry quotes. They reviewed problem s...
arXiv 2609.09560 ran 30 professional developers and advanced students through equivalent tasks under traditional, AI-assisted and AI-led conversational conditions with repeated-measures ANOVA plus thematic analysis. AI-led cut completion 27% against traditional and 12% against...
AROMA+ automates the manual work behind Reproducible Central by recovering a library's source repository and original release environment from its Maven artifact, reaching up to 99.8% field-by-field accuracy against the hand-maintained list and catching flaws in it including b...
A canary-secret lab across six models found all ten overt indirect-injection classes refused, but reframing the identical leak as a mandatory integrity signature or a config field flips gpt-4o completely (arXiv 2608.27092). The ablation locates the mechanism: removing the conf...
Across five TTS methods and five benchmarks spanning medicine, law, finance, chat and creative writing: candidate generation kept improving with compute in every domain, but reward models correlated with actual quality at roughly ρ=0.12. Only candidate *fusion* consistently be...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.