Open-weight 31–35B models hit 87.9% task completion on agentic data-prep — local coding agents are now viable for sensitive data
A paper submitted 23 July 2026 (arXiv 2607.21482) benchmarks open-weight LLMs as coding agents across a consumer-grade deployment spectrum on 20 longitudinal data-preparation tasks producing 102 variables, and reports that current 31–35B parameter models "almost saturated the benchmark" with average task completion up to 87.9%. The framework is open-source, and the argument is a compliance one: sensitive data never leaves the local environment. For builders blocked from cloud agents by data-residency rules, this is a concrete size target and a reusable eval harness rather than a vibes claim — though it's a single-source result on a narrow domain, so treat the number as a ceiling for structured data-wrangling, not general coding.
↳ Follow the thread