LightOnOCR 3 finally makes OCR output point back at the page LightOn released LightOnOCR 3, an Apache 2.0 OCR model family in 0.8B, 1B and 4B sizes that returns a bounding box for every extracted block so chart values and table cells can be traced back to their exact location on the scanned page. The 4B model is the recommended variant, though grounding adds roughly 25 percent more tokens than plain transcription and the vLLM setup is tightly pinned, with the 1B model breaking if Transformers is upgraded independently. LightOn's 75.1 ParseBench score is its own figure, and the public leaderboard places KDL Frontier Parser nano slightly above it. LightOnOCR 3 dropped and the part I care about is not the benchmark number. It is grounding. Ask for it and every extracted block comes back with a bounding box, so a chart value or a table cell can point back to the exact spot on the page you scanned. That is the missing piece for anyone stuffing PDFs into a RAG pipeline and hoping the numbers are right. Three sizes under Apache 2.0, 0.8B, 1B and 4B, with 4B as the one LightOn recommends. Two catches worth knowing. Grounding spits out about 25 percent more tokens than plain transcription, and the vLLM setup is pinned hard, so upgrade Transformers on your own and the 1B model breaks. Also the 75.1 ParseBench score is LightOn's own number, and the public leaderboard has KDL Frontier Parser nano slightly above it. Test it on your ugliest scans before you trust it. We wrote this up at ByteForward with the model choices, prompt modes and an eval checklist