Fetching from the wire…
Public story · 2026-03-19 · source-backed
ArXiv 2603.17946 enables upgrading models from grouped-query attention (GQA) to multi-head latent attention (MLA) via covariance-aware rank-enhanced decomposition. By preserving covariance structure during low-rank decomposition, CARE retains quality while gaining MLA's KV-cache efficiency. Practitioners holding GQA-based models (LLaMA, Qwen lineage) can now retrofit MLA compression as a post-training operation. arXiv
Each link below shares sources, entities, or timing with this story.
Meta released Llama / Shared entities / What happened next
Linked by a graph relationship (Meta released Llama); both cover Llama, Qwen; picks up the Llama thread on 2026-08-16.
Linked by a graph relationship (Meta released Llama); both cover Llama, Qwen; picks up the Llama thread on 2026-05-01.
Linked by a graph relationship (Meta released Llama); both cover Llama, Qwen; picks up the Llama thread on 2026-04-02.
Qwen benchmarked against Claude / Shared entity: Qwen / Shared topic / What happened next
Linked by a graph relationship (Qwen benchmarked against Claude); both cover Qwen; overlapping topics (attention, model).
Meta released Llama / Shared entity: Qwen / What happened next / Tension
Linked by a graph relationship (Meta released Llama); both cover Qwen; picks up the Qwen thread on 2026-08-09.
Alibaba released Qwen / Shared entity: Qwen / What happened next / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover Qwen; picks up the Qwen thread on 2026-06-28.
Ollama supports Qwen / Shared entity: Qwen / What happened next / Tension
Linked by a graph relationship (Ollama supports Qwen); both cover Qwen; picks up the Qwen thread on 2026-04-23.
Alibaba released Qwen / Shared entity: Qwen / What happened next / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover Qwen; picks up the Qwen thread on 2026-04-21.