GB/T-Bench Puts the Human-LLM Gap in Rule-Intensive Document Review at 0.328 vs 0.664, Halved by Multi-Agent Skill Coordination
GB/T-Bench (arXiv 2608.06312, Aug 6) is the first benchmark for structured review of national standard documents, using China's GB/T standards as a testbed with a taxonomy of 25 diagnosable error types across structure, scope alignment, normative modality, terminology consistency, and normative references. A controllable counterexample generator turned 488 documents into 7,306 traceable review error instances, scored under a protocol requiring exact matches on error location, dimension, and type. Across 14 mainstream LLMs the strongest scored 0.3280 CMCS against 0.6640 for human experts; their GB/T-Reviewer multi-agent framework — global inspection, targeted diagnosis, rule scanning, result verification — lifts the best to 0.5094.
↳ Follow the thread