Research
DGF-Bench: Gemini 3.8 Flash Passes 94.98% of Governance Gates, but Only 76.92% of Projects Clear Every Gate
Jeremy Canale's arXiv 2609.29345 treats each enterprise governance review gate as an executable contract and tests agents on 300 synthetic projects with 899 evaluable runs. Strict gate success was 94.98% for Gemini 3.8 Flash, 83.29% for GPT-5.6 Luna and 74.18% for DeepSeek v4.1 Flash, but complete-route success fell to 76.92%, 42.33% and 24.67%. A deterministic control passed all 1,700 gates given the rules and structured facts, which puts the failure in execution rather than judgment. The paper also derives a residual-work threshold showing that automating most cases can still increase total human labor.
Source
↳ Follow the thread