Reddit
Prefill-as-a-Service: Kimi/Moonshot Proposes Cross-Datacenter KVCache Transfer Architecture — 54% Throughput Gain
A new arXiv paper from Moonshot AI proposes Prefill-as-a-Service (PrfaaS), an architecture that offloads long-context prefill computation to dedicated clusters and transfers resulting KVCache across datacenters. The key enabler is hybrid-attention architectures like Kimi Delta Attention (KDA) that reduce KVCache growth by an order of magnitude vs dense attention. In testing with a 1T-parameter hybrid model, PrfaaS achieved 54% higher throughput than homogeneous PD baselines and 32% over naive heterogeneous setups. This could fundamentally change how large-model inference is deployed at scale.
Source
↳ Follow the thread