datalab-to/chandra: State-of-Art OCR Model for Complex Tables, Forms, and Handwriting — 5,758 Stars (+546/day)
GitHub·medium signal
Chandra OCR 2 is a 4B parameter model scoring 85.9% on the olmOCR benchmark (state of the art) with support for 90 languages. Handles complex tables with merged cells, reconstructs form checkboxes and radio buttons, reads cursive handwriting, and preserves multi-column layouts. Outputs structured HTML/Markdown/JSON. Apache 2.0 licensed, free for personal use and startups under $2M revenue.