Fetching from the wire…
Models2026-07-12 · source-backed
In his July 9 annotated write-up he flags Sol at Medium reasoning as a likely upgrade over 5.5 xhigh for coding, and calls the programmatic-tool-calling and multi-agent API additions the genuinely interesting parts (simonwillison.net). Sol also sets a new high of 53.6 on Agents' Last Exam, beating Claude Fable 5 by 13.1 points at roughly a quarter of the estimated cost. His real complaint: figuring out which variant times which reasoning-effort to use is now the hardest part of adopting the family. He's not wrong. Three tiers times four effort levels is twelve combos to reason about per task.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Simon Willison released LLM); both cover Claude Fable, GPT, July, Simon Willison; reported by the same outlet (simonwillison.net).
Claude Code benchmarked against GPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover Claude Fable, GPT, Simon Willison; reported by the same outlet (simonwillison.net).
Claude Code benchmarked against GPT / Shared entities / Same source domain / What happened next / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, July, Simon Willison; reported by the same outlet (simonwillison.net).
Claude Code benchmarked against GPT / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover Claude Fable, Simon Willison; reported by the same outlet (simonwillison.net).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, Simon Willison; reported by the same outlet (simonwillison.net).
Simon Willison uses Fable / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Simon Willison uses Fable); both cover Claude Fable, GPT, July; overlapping topics (claude, fable).
Claude Code benchmarked against GPT / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, July; overlapping topics (agent, claude, fable).
Claude Code benchmarked against GPT / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, July; overlapping topics (agent, claude, cost).