OpenAI field report: eight agent-assisted scientific computing projects, and agents can't judge scientific validity
OpenAI published a field report examining eight agent-assisted scientific computing projects, mostly in life sciences, where researchers used Codex and Claude Code for everything from packaging updates and targeted optimizations to full language migrations and GPU-native redesigns. The core finding is a division of labor: agents materially accelerated development but could not reliably judge scientific validity, requiring human verification at every stage, and researchers shifted from implementation to orchestration — defining goals, chunking work, validating outputs. The report flags long-term stewardship as unsolved, warning that AI-assisted rewrites become tomorrow's abandoned code without clear ownership.
Source
↳ Follow the thread