Micro Language Models Enable Sub-Second Responses on Smartwatches and Glasses Without Cloud Roundtrip
arXiv·high signal
Cheng and Chen address the gap where edge devices (smartwatches, smart glasses) cannot continuously run even 100M-1B parameter models due to power constraints, while cloud inference introduces multi-second latencies. Their micro language model architecture targets sub-100M parameter footprints that run entirely on-device with instant response times. This extends the sub-billion model trend (Gemma 3 270M, SmolLM2 135M) to wearable-class hardware with active power budgets under 1W.