Fetching from the wire…
Public story · 2026-08-24 · high
BC-Bench evaluated 101 bug-fix tasks from two production Business Central repos written in AL, a DSL with little public training data.
Why now: The paper posted August 24, 2026.
Microsoft researchers built a 101-task bug-fixing benchmark from two production repos written in AL, the DSL behind Dynamics 365 Business Central. They adapted SWE-bench's method for a language with almost no public training data to learn from.
The results cut against a growing assumption that agent harness design matters more than model choice. In AL bug-fixing, gaps between frontier models exceeded gaps between the two harnesses tested, per the BC-Bench paper.
Gains models showed on general-purpose coding benchmarks didn't consistently carry over to AL bug-fixing either, the study found.
That's a problem beyond Business Central shops. Teams building on any DSL with thin public training data can't assume a model's benchmark rank holds for their codebase. For teams evaluating AI coding assistants, the difference matters. A model topping general leaderboards isn't guaranteed to be the best pick for a legacy or niche-language codebase.
The paper doesn't say why the harness effect shrank on AL specifically, only that it did in this test.
Each link below shares sources, entities, or timing with this story.
Microsoft criticizes Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Bench, SWE; reported by the same outlet (arxiv.org).
Microsoft supports MCP / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft supports MCP); both cover Bench, SWE; overlapping topics (agent, between, code, harness, model).
Microsoft criticizes Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft criticizes Claude Code); both cover Bench, SWE; overlapping topics (behind, benchmark, between, code, model).
Microsoft competes with OpenAI / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Microsoft competes with OpenAI); both cover Microsoft, SWE; overlapping topics (benchmark, code, model).
Anthropic partners with Microsoft / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic partners with Microsoft); both cover Bench, SWE; overlapping topics (agent, benchmark, model).
Microsoft released OpenForgeRL / Shared entity: Microsoft / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Microsoft released OpenForgeRL); both cover Microsoft; reported by the same outlet (arxiv.org).
Microsoft competes with OpenAI / Shared entity: Bench / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft competes with OpenAI); both cover Bench; reported by the same outlet (arxiv.org).
Microsoft released OpenClaw / Shared entity: SWE / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Microsoft released OpenClaw); both cover SWE; reported by the same outlet (arxiv.org).