Fetching from the wire…
Top 5 · 2026-06-15 · source-backed
OpenRouter released Fusion, a compound API that fans each prompt out to a panel of models, synthesizes their answers, and returns one response (OpenRouter). On Perplexity's DRACO deep-research benchmark, 100 tasks across 10 domains, a Fable 5 + GPT-5.5 fusion scored 69.0% versus 65.3% for Fable 5 alone. The number I keep rereading: a budget panel of Gemini 3 Flash + Kimi K2.6 + DeepSeek V4 Pro beat solo GPT-5.5 and Opus 4.8 at about half Fable 5's price.
So a committee of cheaper models, voting and synthesizing, beats a single expensive model. That's a genuinely different way to think about the stack. We've spent two years assuming the move is "pick the best single model and route everything to it." Fusion says the best answer might be an ensemble where no individual member is frontier-class.
The catch is real and they're upfront about it: calls run 2-3x longer. For deep research, agentic background tasks, anything where a human isn't tapping their foot waiting, that latency is free. For interactive chat or anything user-facing in a tight loop, it's a dealbreaker. So this isn't a universal upgrade. It's a tool for a specific shape of problem, and the shape is "quality matters more than speed and I'm cost-sensitive."
Notice how this rhymes with the DeepSeek story. Both are arguments that you no longer have to pay frontier prices for frontier-adjacent quality, you just have to be willing to architect for it. Cheap open weights on one side, cheap ensembles on the other, and the single-frontier-model default getting squeezed from both directions.
What I'd do: if you're running deep-research or batch-analysis workloads, benchmark Fusion's budget panel against whatever single model you're paying for now, on your own eval set. If it holds quality at half the cost and you can eat the latency, that's found money. I'm skeptical of the universal framing though. "Beats frontier" on one benchmark in one domain category is not "beats frontier" everywhere, and I'd want to see it on coding and tool-use tasks before I believe the headline.
Each link below shares sources, entities, or timing with this story.
DeepSeek released deepseek-v4-flash / Shared entities / Shared topic / What happened next
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover DeepSeek, Gemini, GPT, Kimi K2; overlapping topics (cost, deepseek, model, task).
Gemini competes with ChatGPT / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Gemini competes with ChatGPT); both cover Fable, Gemini, GPT, Kimi K2; overlapping topics (best, model).
Gemini competes with ChatGPT / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with ChatGPT); both cover Flash, Gemini, GPT, Opus; overlapping topics (cost, frontier, model, task).
Opus built by Anthropic / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Opus built by Anthropic); both cover DeepSeek, Fable, GPT, Opus; overlapping topics (cost, deepseek, fable, model).
Linked by a graph relationship (Opus built by Anthropic); both cover DeepSeek, Fable, GPT, Opus; overlapping topics (beat, benchmark, fable, model).
Kilo Code uses OpenRouter / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kilo Code uses OpenRouter); both cover DeepSeek, Gemini, GPT, Opus; overlapping topics (benchmark, frontier, model, quality).
DeepSeek released DeepSeek V4 / Shared entities / Shared topic / Tension
Linked by a graph relationship (DeepSeek released DeepSeek V4); both cover DeepSeek, Fable, Flash, GPT; overlapping topics (cost, deepseek, frontier, model).
OpenHands uses GPT / Shared entities / Shared topic / What happened next
Linked by a graph relationship (OpenHands uses GPT); both cover DeepSeek, Fable, Flash, GPT; overlapping topics (cost, model, task).