Fetching from the wire…
Infra2026-08-11 · source-backed
Mike Ayles' August 10 writeup runs an INT4 transformer, about 1.5 MB of weights trained on TinyStories, entirely in on-chip SRAM on a Xilinx Kria KV260 with zero DDR in the token-generation loop. Peak aggregate throughput is 59,965 tok/s at 200 MHz, with the live demo sustaining ~21,300 tok/s under 2,000 concurrent connections. Same board's ARM A53 cores manage 11 tok/s; a laptop RTX 3050 Ti manages 719. The entire result is trading 20 GB/s shared DDR for hundreds of GB/s of on-chip bandwidth.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover August, Same; earlier August coverage from 2026-08-09.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM, Same; earlier LLM coverage from 2026-07-31.
Linked by a graph relationship (Simon Willison released LLM); both cover August, LLM; earlier August coverage from 2026-08-07.
Linked by a graph relationship (Simon Willison released LLM); both cover August, LLM; earlier August coverage from 2026-08-05.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
Simon Willison released LLM / Shared entity: Peak / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover Peak; earlier Peak coverage from 2026-05-09.