Agents
A controlled ablation on a production science agent finds the model choice dominates topology and prompting, and a PPO policy nearly matches it for free
arXiv 2608.25215 (25 Aug) ablates federation topology, RL versus LLM harness, model, and prompt expertise on a verifiable protein-function characterization task routed across tools. Model choice dominated everything else, Opus at roughly 92-94% against o4-mini at 40-50%, while federation across institutional boundaries imposed a negligible performance penalty. The result builders should note: a PPO policy hit 88% at zero token cost with the fastest latency and perfect consistency but no reasoning trace, so for routine verifiable tasks a cheap deterministic policy is close to frontier, and prompt dependence was largest exactly when the task was hardest.
Source
↳ Follow the thread