Agents
Faraday: a 27B agent trained specifically to replicate research beats Claude Opus 4.8 and GPT-5.5 on held-out replication
arXiv 2608.13331 (2026-08-13) trains a 27B-parameter 'AI Scientist' agent on the narrow task of reproducing published research, and reports it surpassing Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. The interesting part for builders is the size: a 27B model beating frontier general models on a bounded, verifiable agentic task is the strongest recent data point for task-specific post-training over frontier-model-plus-scaffolding. Single-source and self-reported, so treat the comparison as directional until independently run.
Source
↳ Follow the thread