'Echo' Claims Fable-Level Output at One-Third the Cost by Routing Across GLM-5.2 and Kimi K2.7
A Show HN posted July 23 pitched Echo as a task-adaptive router over a pool of open-weight models — GLM-5.2, Kimi K2.7 and others — allocating "the intelligence it needs" per task rather than committing to one endpoint, and reports matching Claude Fable on the evaluated task mix at roughly one-third total inference cost while outperforming every individual open-weight model tested. It drew several hundred points and a couple hundred comments, with the dominant critique being reproducibility: commenters want the published task mix and side-by-side outputs, not aggregate claims. Treat the numbers as unverified vendor benchmarks, but the pattern — routing as the cost lever now that open weights are close on quality — is the thing worth tracking.
↳ Follow the thread