Agents
EEBench scores agents on real circuit boards with SPICE, and Claude Opus 5 leads at 61.6% while GPT-5.6 Sol trails at 39.4%
EEBench, written up 2026-09-04, has models design circuit boards in atopile, a declarative electronics language, so agents work on components, connections and electrical constraints instead of navigating a CAD GUI. Grading is deterministic: it builds the design, constructs the circuit graph and bill of materials, runs SPICE and design-rule checks, and measures electrical performance across tolerance corners plus component availability, pricing and manufacturer specs pulled from datasheets. On the 2026-09-01 board Claude Opus 5 leads at 61.6%, followed by Grok 4.6 at 57.1%, Claude Fable 5.1 at 56.4%, Fable 5 at 54.3% and Opus 4.8 Max at 51.4%, with GPT-5.5 at 42.3% and GPT-5.6 Sol at 39.4%.
Source
↳ Follow the thread