Fetching from the wire…
Infra2026-07-26 · source-backed
Tsinghua's Chen, Piao, Feng, Hu, and Li convert PoT programs into parameterized cache objects reusable across structurally similar requests, with one small model doing double duty: semantic variable extraction on cache hits and speculative drafting during target-LLM generation (arXiv 2607.20507). Up to 3.1x lower latency and 2.8x higher throughput under parallel serving on WebShop, Formula, and CodeTAT-QA. The architectural framing is the takeaway: small models earn their keep as lightweight interface layers around a cache, not as cheap substitutes for the big model.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-14.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-22.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-10.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-10.