Reddit
Practitioner Clocks Liquid AI's LFM2.5-2.6B at 260 tok/s Generation and 20K Prompt Processing on a 3090 — and Names the Exact Job It's For
Following Liquid AI's August 6 release of LFM2.5-2.6B (128K context, tool calling, open weights, ~2.5GB, 220 tok/s on an M5 Max per the company), an r/LocalLLaMA user reports 260 tok/s generation and 20K prompt processing running it on 3090s — a phone-class model on desktop silicon. Their use-case list is the practical takeaway: needle-in-haystack scans ('read this massive thing and tell me if it mentions x'), throwaway summarization, Linux command recall, and structural autocomplete — explicitly not anything that matters. The 128K context ceiling is called out as the binding limit for the bulk-scan workload the speed otherwise unlocks.
Source
↳ Follow the thread