Fetching from the wire…
Skills2026-07-30 · source-backed
A July 28 study on HumanEval+, MBPP+, and LiveCodeBench found real original tests moved Qwen3.6 on LiveCodeBench from 13.1% to 39.4%, while stronger-model-generated synthetic tests added 1.7 points at p = .701, statistically indistinguishable from nothing. Spend your retrieval budget finding actual tests in the repo, not manufacturing plausible ones.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Earlier coverage / Tension
Both cover July, Spend; reported by the same outlet (arxiv.org); earlier July coverage from 2026-07-21.
Shared entities / Same source domain / Tension
Both cover July, MBPP; reported by the same outlet (arxiv.org); pushes against this story (versus).
Shared entities / Same source domain / Earlier coverage
Both cover HumanEval, July; reported by the same outlet (arxiv.org); earlier HumanEval coverage from 2026-07-28.
Shared entity: July / Same source domain / Shared topic / Earlier coverage
Both cover July; reported by the same outlet (arxiv.org); overlapping topics (finding, test).
Shared entity: Qwen3 / Same source domain / Shared topic / Earlier coverage
Both cover Qwen3; reported by the same outlet (arxiv.org); overlapping topics (generating, instead).
Shared entities / Earlier coverage / Tension
Both cover LiveCodeBench, Qwen3; earlier LiveCodeBench coverage from 2026-03-22; pushes against this story (vs).
Shared entity: July / Same source domain / Earlier coverage / Tension
Both cover July; reported by the same outlet (arxiv.org); earlier July coverage from 2026-07-28.
Both cover July; reported by the same outlet (arxiv.org); earlier July coverage from 2026-07-27.