Fetching from the wire…
Research2026-08-06 · source-backed
arXiv 2608.04975 is a domain-expert audit of all 65 SciCode problems, a component of the Artificial Analysis Intelligence Index and a standing eval in government and national-lab suites. 192 of the 263 defects, spread across 91% of main problems, wrongly rejected correct solutions via non-reproducible gold answers, over-tight tolerances, or self-contradictory specs. 78% of score-suppressing defects required specialized physics or math to spot, not proofreading. Re-evaluating twelve frontier snapshots on the corrected set lifts subproblem accuracy from 45-60% to 84-98% and main-problem accuracy from 9-27% to 69-92%. The widely-cited 2026 scientific-coding plateau was the instrument.
Each link below shares sources, entities, or timing with this story.
The Wiggle Framework stress-tested 9 frontier models across 14 judging tasks. The damning part: flips were almost always net-corrupting relative to ground truth. Pressure moved judges away from the right answer, not toward it. arXiv 2608.12645 If you use LLM-as-judge anywhere...
SpecPath found 35 of 100 passing implementations broke when only the revision path changed, with aggregate accuracy looking identical across paths. Build your eval set from real multi-turn clarification threads with amendments and reversals. Path sensitivity is invisible to st...
Auditing Qwen2.5-7B-Instruct on RGB and HotpotQA with a hallucination detector, NLI entailment and an LLM judge, INT8 is near-lossless on accuracy and faithfulness (arXiv 2608.30996). INT4 lowers accuracy, and among answers that stay factually correct, over 90% of faithfulness...
Per Willison, 1.57T total with 48B active per token, open weights already mirrored as GGUF, landing at 53 on the Artificial Analysis Intelligence Index with LiveCodeBench 93.50, MMLU Pro 87.50 and SWE-bench Verified 80.60. Independent-evaluation trackers showed zero third-part...
Reported August 7, Grok 4.6 reuses the 1.5T-parameter V9 foundation from 4.5, with gains attributed to improved SFT and RL rather than scale. xAI has published no benchmarks, model card, or scores, so every circulating number is speculation until independent Arena results land...
Xiaomi released MiMo-V2-Pro, a 1T parameter open-weight model with 1M token context and native agent-orchestration capabilities optimized for tool-calling and multi-step reasoning. Available on OpenRouter at ~67% cheaper than Claude Sonnet 4.6, ranking #8 on the Artificial Ana...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.