Fetching from the wire…
Infra2026-07-26 · source-backed
Ponnusamy, Sahni, Wang, and Tri Dao attack a real serving problem: existing sampling implementations accelerate only parts of the logit-processing/token-selection/verification pipeline, need multiple kernel launches, or assume every request in a batch samples identically, which breaks CUDA Graph execution for dynamic workloads (arXiv 2607.20475). Their kernel suite handles grammar-constrained decoding, penalties, filtering, and speculative verification in a single batched kernel while staying CUDA Graph compatible, with hierarchical two-stage top-k giving up to 10x, and up to 16x across heterogeneous workloads. Drop-in win if you run mixed sampling configs behind one endpoint.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-14.
Simon Willison released LLM / Shared entity: Drop / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Drop; earlier Drop coverage from 2026-07-01.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-22.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-10.