Agents
BC-Bench finds model choice matters more than agent harness on ERP domain-specific code
Microsoft researchers released BC-Bench, 101 manually curated tasks pulled from two Microsoft-owned production repositories for AL, the DSL behind Dynamics 365 Business Central. It adapts the SWE-bench method to an ecosystem with scarce public training data and heavy environment provisioning, and it scores test generation and multimodal problem statements, not just functional code. In the bug-fixing category, differences between frontier models exceeded differences between the two agent harnesses, and gains reported on general-purpose benchmarks did not consistently transfer to AL.
Source
↳ Follow the thread