Fetching from the wire…
Research2026-06-18 · source-backed
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0% on Last-Exam tasks. The next time someone says agents are "job-ready," this is the number to quote back. (arXiv 2606.05405)
Each link below shares sources, entities, or timing with this story.
Claude benchmarked against Codex / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (Claude benchmarked against Codex); both cover Bench, Codex, GPT, Terminal; overlapping topics (agent, codex, gpt-5).
OpenAI released Codex / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Codex); both cover Bench, Codex, GPT, Terminal; overlapping topics (gpt-5, task).
Codex competes with Claude Code / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex, GPT; reported by the same outlet (arxiv.org).
OpenAI released Codex / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Codex); both cover Bench, Codex, GPT, Terminal; overlapping topics (codex, gpt-5).
Codex released Windows / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Codex released Windows); both cover Bench, Codex, GPT, Terminal; overlapping topics (codex, gpt-5).
Cursor benchmarked against Codex / Shared entities / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (Cursor benchmarked against Codex); both cover Codex, GPT; reported by the same outlet (arxiv.org).
JetBrains supports Codex / Shared entities / Shared topic / What happened next
Linked by a graph relationship (JetBrains supports Codex); both cover Codex, GPT, Last Exam; overlapping topics (agent, codex, exam).
Claude benchmarked against Codex / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Claude benchmarked against Codex); both cover Bench, GPT, Terminal; overlapping topics (agent, codex, tier).