Reddit
A One-Image Analog Meter Test Is Being Used as r/LocalLLaMA's Informal Vision Benchmark — Ground Truth 37461
A r/LocalLLaMA post (147 upvotes, 60 comments) asks readers to run a single utility-meter photo through their preferred local vision model and report the digits; the correct reading is 37461. The 0.41 comment-to-score ratio indicates heavy result-sharing rather than passive upvoting, making the thread a crowd-sourced VLM comparison across whatever people actually run locally. It is a useful reminder that the tasks practitioners care about — reading a slightly skewed real-world dial — are not what published OCR benchmarks measure, and that per-model failure here is easy to reproduce yourself in seconds.
Source
↳ Follow the thread