Fetching from the wire…
Security2026-07-26 · source-backed
Majidi, Mireshghallah, and Taram demonstrate the first attacks inferring proprietary model and deployment details from per-token generation timing over a remote API (arXiv 2607.20723). One attack detects whether a provider runs speculative decoding and recovers the draft model's context length, measuring Google Gemini Flash 2.5 at roughly a 128K-token draft window. The other recovers layer count, hidden dimension, and attention-head count by modeling latency on NVIDIA GPUs, landing the near-correct Llama configuration in the top-10 more than 90% of the time. Streaming APIs are a side channel. Nothing about the response content has to leak.
Each link below shares sources, entities, or timing with this story.
Meta released Llama / Shared entity: Llama / Shared topic / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover Llama; overlapping topics (closed, context, model).
Meta released Llama / Same source domain / Shared topic
Linked by a graph relationship (Meta released Llama); reported by the same outlet (arxiv.org); overlapping topics (context, model).
Meta released Llama / Shared entity: Llama / Earlier coverage
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-07-19.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-06-15.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-05-02.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-04-02.
Linked by a graph relationship (Meta released Llama); both cover Llama; earlier Llama coverage from 2026-03-12.
Meta released Llama / Shared topic
Linked by a graph relationship (Meta released Llama); overlapping topics (closed, model).