Fetching from the wire…
Skills2026-07-30 · source-backed
A July 28 study on HumanEval+, MBPP+, and LiveCodeBench found real original tests moved Qwen3.6 on LiveCodeBench from 13.1% to 39.4%, while stronger-model-generated synthetic tests added 1.7 points at p = .701, statistically indistinguishable from nothing. Spend your retrieval budget finding actual tests in the repo, not manufacturing plausible ones.
Each link below shares sources, entities, or timing with this story.
Ockhamareto (arXiv 2608.24473) reinforces a unit-test rollout only when it's non-dominated on both mutation-killing and test count, then ties each test's killing power back to specific source tokens. Against MIST-RL that's a 3.4x better per-test trade-off, plus 30 to 35 percen...
A stage-wise study of self-refinement across 5 benchmarks with 6 sizes of Qwen3 and 4 sizes of Gemma 3 found larger generators and refiners generally improve the pipeline, and an undersized refiner can actively hurt, but results are highly insensitive to critic size. Including...
PR #26062, "server: support MCP stdio," by ngxson, merged into ggml-org/llama.cpp on July 25 (r/LocalLLaMA). It landed alongside #26061 (vendored subprocess.h, merged July 24) and pwilkin's #26075 integration-and-tests PR. Until now, llama-server's web UI could only talk to MC...
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Z.ai's 744B MoE is being ranked as July's strongest open-source model at a reported 91.2% GPQA Diamond and 62.1% SWE-bench Pro, at a fraction of frontier API pricing, and it's listed in Ollama's supported-model line. That's the distinction that matters this week: Kimi K3's wei...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.