Fetching from the wire…
Public story · 2026-08-07 · high
A new BFCL v4 benchmark shows GPT-5.6 gaining 10.6 percent when it writes code to call tools instead of emitting JSON.
Why now: The BFCL v4 numbers give framework maintainers a controlled comparison instead of a guess, landing while most agent frameworks still default to JSON tool calling.
A model that writes code to call tools outperformed one filling out a JSON schema in 11 of 14 models tested on BFCL v4, per arXiv preprint 2608.06370.
That's a problem for teams that built their agent loop around JSON-schema tool calls as the default. The GPT-5.6 family gained 10.6% in accuracy under programmatic tool calling, and the gain held up under parallel execution in 13 of 14 models.
Under context degradation, the JSON baseline's accuracy fell 2.3% on average. Programmatic calling held steadier.
Yes, three of the 14 models still did better with JSON tool calling, so this isn't a clean sweep.
The advantage tracked model capability across release generations instead of fading. A one-off quirk in a single model family would shrink going backward through older, smaller models. This one didn't.
Framework maintainers who keep JSON as their only default are tuning for the weakest model in their fleet, not the one running in production. Worth watching whether the three holdout models share an architecture or training recipe. That distinction would show whether this is a permanent split or a one-time transition.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source / Shared topic
Both cover BFCL, GPT, JSON, PTC; cite the same source (arXiv 2608.06370); overlapping topics (agent, calling, json, model, tool).
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
GPT competes with Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT, JSON; overlapping topics (json, model).
GPT competes with Claude / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (GPT competes with Claude); both cover GPT; reported by the same outlet (arxiv.org).