Fetching from the wire…
Research2026-08-30 · source-backed
Most visual token pruning runs after the encoder, leaving encoder latency untouched. PACE is training-free in two stages: an Adaptive Pixel Compressor scores visual information density before encoding and downsamples redundant input, then a Dynamic Dual-Attention Extractor keeps tokens using both internal encoder signals and semantic signals from the LLM. 3.1x speedup in time to first token, code at github.com/jjL357/PACE. (arXiv 2608.27206)
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM, Most; earlier LLM coverage from 2026-07-31.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-16.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-27.
Simon Willison released LLM / Shared entity: Most / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Most; earlier Most coverage from 2026-08-24.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-08-21.