Reddit
Liquid AI's LFM2.5-VL-3B Jumps RefCOCO Grounding Precision From 57.1 to 87.9 and Decodes 20 tok/s on a Galaxy S26 Ultra
Liquid AI released LFM2.5-VL-3B on August 12, a 3.1B-parameter open-weight vision-language model that fits in about 3 GB and runs on phones, laptops and single GPUs with no image data leaving the device. It averages 80.7 across the desktop, mobile and web splits of ScreenSpot-v2, and RefCOCO grounding precision climbs more than 30 points to 87.9. Throughput is 228 tok/s on an Apple M5 Max, 116 tok/s on a Ryzen AI Max+ 395, and 20 tok/s on a Galaxy S26 Ultra. Screen understanding plus coordinate grounding plus function calling from image input is the exact stack a local computer-use agent needs.
Source
↳ Follow the thread