News
OpenAI Field Report: Coding Agents Hit 60x Speedups Rewriting Research Software — and Can't Tell When They're Wrong
OpenAI and academic partners published eight real deployments where coding agents modernized neglected scientific software, rewriting roughly 20,000 lines of legacy C++ into Rust. RustQC collapsed 15 separate QC tools into one program, cutting runtime from 15h34m to 14m54s (>60x); HelixForge's GPU rewrite of a synthetic genomics generator ran 59.6x faster. Five projects used Codex alone and three combined Codex with Claude Code — but the report's honest finding is that agents executed well-defined tasks fast while confidently presenting scientifically wrong results, with every win resting on a human defining correctness and building the verification machinery.
Source
↳ Follow the thread