Research
ContractScrub: Only One Frontier Model Reaches 0.75 Macro Recall on Lawyer-Written Contract Errors
arXiv 2608.20204 introduces the first benchmark for contract scrubbing, the final review pass that hunts misused defined terms, incorrect cross-references, and inconsistent language in transactional agreements, built from contracts hand-crafted by experienced lawyers across those error categories. Despite the task looking like a natural fit for long-context reasoning, consistency checking, and NER, frontier models perform poorly and only one reaches 0.75 macro average recall. The authors read this as evidence that strong scores on general benchmarks do not transfer to narrow high-value professional work, and that domain-specific evals are the only way to measure real-world exposure.
↳ Follow the thread