Fetching from the wire…
Public story · 2026-08-05 · high
The paper's own reasoning traces show the models spot the odd element, then override it with what the pattern predicts.
Why now: The paper's numbers surfaced in the August 5, 2026 coverage, worth weighing before trusting screenshot-to-code output.
A benchmark hides one altered measurement in a repeated UI pattern and asks five multimodal models to catch it, per arXiv paper 2608.03691.
Teams feeding screenshots into code-generation tools are trusting these models to reproduce what's actually rendered, not what a layout predicts. Mean recovery accuracy across the five models lands at 21.17% for card width and 7.89% for text font size, per the paper.
Results split hard by model. The paper names two of the five directly: Codex-5.3 recovers card width best at 68.61% but collapses to 13.89% on font size. Flash-3.0 shows a 96.11% bias toward the pattern's predicted font size, reporting what the layout implies almost every time regardless of what's actually on screen.
The reasoning traces are the sharper finding. The paper shows models correctly naming the anomalous element mid-thought, then overriding themselves and reporting the pattern-consistent answer anyway. The model sees the outlier and reports the average.
That failure mode won't throw an error. A layout will look right and measure wrong. A card ends up a few pixels narrower than spec, a caption a size off, both silently normalized back to what the pattern expects.
Each link below shares sources, entities, or timing with this story.
Codex competes with Claude Code / Shared entity: Codex / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex; reported by the same outlet (arxiv.org).
OpenAI released Codex / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Codex); both cover Codex, Flash; earlier Codex coverage from 2026-04-25.
Codex competes with Claude Code / Shared entity: Codex / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex; reported by the same outlet (arxiv.org).
Codex uses ChatGPT / Shared entity: Codex / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Codex uses ChatGPT); both cover Codex; reported by the same outlet (arxiv.org).
Codex competes with Gemini / Shared entity: Flash / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Gemini); both cover Flash; reported by the same outlet (arxiv.org).
Codex competes with Claude Code / Shared entity: Codex / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex; reported by the same outlet (arxiv.org).
Port22 uses Codex / Shared entities / Earlier coverage
Linked by a graph relationship (Port22 uses Codex); both cover Codex, Flash; earlier Codex coverage from 2026-08-02.
Codex competes with Gemini / Shared entities / Earlier coverage
Linked by a graph relationship (Codex competes with Gemini); both cover Codex, Flash; earlier Codex coverage from 2026-07-31.