Fetching from the wire…
Public story · 2026-08-26 · high
A replication across three model pairs finds GPT-4o barely gains from prompting tricks that still help Qwen2.5, Mistral mixed.
Why now: The paper posted in August 2026 and tests three model upgrade pairs: GPT-3.5 to GPT-4o, Qwen2 to Qwen2.5, and Mistral-7B-Instruct to Mistral-Large.
A replication study finds GPT-4o gains little from chain-of-thought and few-shot prompting that still boosts Qwen2.5 7B, per a paper posted to arXiv.
The study ran 19,620 generations across 218 context-rich Python functions, testing Zero-Shot, Few-Shot, Chain-of-Thought, Contrastive CoT, and an adapted Program-of-Thought. It ran those five techniques across three model-version pairs: GPT-3.5-Turbo to GPT-4o, Qwen2 7B to Qwen2.5 7B, and Mistral-7B-Instruct to Mistral-Large. Teams carrying the same prompt templates across a model swap can't assume the old lift still applies.
GPT-4o shows diminishing or negative gains from the same structured prompting that reliably boosted GPT-3.5-Turbo, the researchers found. Their read is that the scaffolding gets absorbed during training. The model already reasons in steps without being told to. Adding chain-of-thought instructions on top adds no lift, and sometimes costs some.
Qwen2.5 7B didn't follow that curve. It still gets a real bump from Few-Shot and Contrastive CoT, the same techniques that stopped paying off on GPT-4o.
Mistral's results split the difference. Mistral-7B-Instruct and Mistral-Large moved in mixed directions across the five techniques, with no clean pattern either way.
Each link below shares sources, entities, or timing with this story.
Qwen benchmarked against Claude / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover GPT, Prompt, Zero; reported by the same outlet (arxiv.org).
Qwen competes with Meta / Shared entities / Earlier coverage
Linked by a graph relationship (Qwen competes with Meta); both cover GPT, Mistral, Qwen; earlier GPT coverage from 2026-05-01.
Qwen competes with Google / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Qwen competes with Google); both cover GPT, Qwen; overlapping topics (becom, model).
Alibaba released Qwen / Shared entities / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover Carrying, Instruct, Qwen, Zero; earlier Carrying coverage from 2026-08-15.
Qwen benchmarked against Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Qwen benchmarked against Claude); both cover GPT; reported by the same outlet (arxiv.org).
Anthropic criticizes Qwen / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Anthropic criticizes Qwen); both cover GPT, Qwen; earlier GPT coverage from 2026-04-21.
Qwen benchmarked against Claude / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Qwen benchmarked against Claude); both cover GPT, Python; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Qwen benchmarked against Claude); both cover GPT, Qwen; reported by the same outlet (arxiv.org).