Science sandboxes catch agents optimizing a metric without ever learning the rules behind it
The framework runs agents through repeated cycles of experimentation, feedback and hypothesis revision against environments spanning wet physical experiments, damp predictive models trained on empirical data, and dry invented rules, so the degree of empirical verifiability becomes a controlled variable. Instantiated on regulatory genomics and protein fitness prediction, frontier agents could be observed hitting quantitative targets while failing to understand the underlying system, and their scientific reasoning degraded specifically when a system's rules fell outside familiar biological priors. That last point is the useful one: apparent scientific competence is partly memorized prior, and the dry-rule arm is what exposes it.
Source
↳ Follow the thread