Fetching from the wire…
Research2026-08-27 · source-backed
Fourteen models including Claude 4.5, GPT-5.2, DeepSeek V4-Pro and the Qwen families run under realistic repository constraints in a containerized framework with file-level, LSP-based and retrieval-based context strategies (arXiv 2608.25939). Invocation Rate is the metric to steal: it checks whether a generated test actually exercises the function it targets, rather than passing vacuously. Results show a substantial gap between standalone and repository-level performance, and richer context does not monotonically improve test reliability.
Each link below shares sources, entities, or timing with this story.
Qwen benchmarked against Claude / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Qwen benchmarked against Claude); both cover CLAUDE, GPT, LSP, PHP; overlapping topics (benchmark, claude, test).
Qwen benchmarked against Claude / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover Claude, GPT, Qwen; reported by the same outlet (arxiv.org).
Qwen benchmarked against Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover Claude, DeepSeek V4, GPT; overlapping topics (claude, deepseek).
DeepSeek released DeepSeek V4 / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (DeepSeek released DeepSeek V4); both cover GPT, Qwen, Rust; overlapping topics (benchmark, context).
Qwen benchmarked against Claude / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover CLAUDE, GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Qwen benchmarked against Claude); both cover CLAUDE, GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Qwen benchmarked against Claude); both cover CLAUDE, GPT; reported by the same outlet (arxiv.org).
Claude benchmarked against Codex / Shared entity: Results / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude benchmarked against Codex); both cover Results; reported by the same outlet (arxiv.org).