A local 4B model drives a full DFT research pipeline at 95.7% extraction precision when deterministic code holds the gates
This audit runs Qwen3:4B open-weights locally through an autonomous materials-science pipeline, extracting parameters from papers, translating them to DFT inputs and driving simulations to convergence under a neurosymbolic split where agents propose and deterministic code disposes. Against 201 expert judgements the extractor reaches 95.7% precision but only 67.3% recall, and full GPU residency matters more than weight or cache precision, raising Matthews correlation from 0.414 to 0.560. Converged results reproduce published lattice constants within 2.3% mean absolute relative error. The side finding is bleak for the field: only 19 of 57 studies audited report enough method parameters to be reproducible at all.
Source
↳ Follow the thread