Content Quality: News article, 829 words (raw count and with link URLs stripped; within the News range of 400-1200). Title is 136 characters (cap 150). Clear structure: Overview, release contents, vendor-reported benchmarks, speed, training notes, and a What We Don't Know section. Appropriately framed as the lightonai authors' own claims.
Source Verification: Snapshots read by gunzip/text extraction; all four exist, are HTTP 200 and contain the full page body (not chrome); sha256 of decompressed content matches the manifest. source-0.html.gz (huggingface.co/blog/lightonai/lightonocr-3, Community Article, 'Published October 8, 2026', authors Said Taghadouini, Adrien Cavailles, Baptiste Aubertin of lightonai): confirms the long Overview quote verbatim ('are now able to output bounding box coordinates with labels for all document regions, image descriptions as well as extract numerical data in figures or charts'), the Apache 2.0 / research and commercial sentence, three sizes (0.8B, 1B, 4B) with the 1B retaining the previous architecture and 0.8B/4B on Qwen3.5, the two modes with 'grounding' as the sole text prompt, 0-1000 normalized coordinates, image and chart block descriptions, all olmOCR-Bench figures (4B 86.3; Infinity Parser Pro 87.6 at 35.1B so 1.3 behind; Chandra 2 85.8 so 0.5 below the 4B; 0.8B 85.5; 1B 84.5), ParseBench (75.1, 74.6, Infinity 74.3, 1B 71.4), fr-bench-pdf2md (74.1, 70.5, 69.6), the LightOnOCR-2-1B asterisk (83.2 excludes headers/footers), the 9-14% fewer output tokens vs Chandra-OCR-2 and Infinity-Parser2-Pro on the same 512 pages, speed section (one H100, vLLM 0.30.0, 400 DPI / 5 MP cap, ~4.8k image tokens, +44% to 4.78 pages/s for 0.8B and +66% to 3.36 pages/s for 4B at 1540 px, 4B 19% faster single page and 21% more pages/s than Chandra-OCR-2 which is also Qwen3.5-4B based), data mixture (2x formulas, 3x tables, 21.9% grounding / 78.1% OCR), 55% retention and ~10% recovery, Qwen3-VL-235B for figure annotation, RL with verifiable rewards after SFT, and the normalization-function passage. The blog's results tables list the 4B model at 4.0B while its own 'Models mentioned' widget shows 0.9B / 1B / 5B for the three models. source-1.html.gz (LightOnOCR-3-4B model card): License apache-2.0 and 'Apache License 2.0'; 'The models come in three sizes'; 'Model size 5B params', BF16 (4,539,265,536 total parameters appears in the page's safetensors metadata); note on Qwen3.5 classes (Transformers >= 5.2, saved with 5.5.4) and thinking disabled as the chat template default; 'grounding' prompt. source-2.html.gz (LightOnOCR-3-0.8B card): Apache 2.0, '0.9B params', same notes. source-3.html.gz (github.com/lightonai/LightOnOCR README): 'A minimal client, CLI and viewer for the LightOnOCR models (versions 1, 2 and 3) served with vLLM, plus the code to reproduce our benchmarks'; 'Grounding ... works with LightOnOCR-3 only. LightOnOCR-1 and LightOnOCR-2 do plain OCR only'; model table 2048 px for 0.8B and 4B, 1540 px for 1B. All claims attributed to the correct source. No fabricated or unsupported specifics found. Minor wording: the article says normalization is applied to 'raw outputs'; the blog text says 'raw inputs' (almost certainly a typo for outputs; the article is an unquoted paraphrase) - not material.
Factual Accuracy: Every number, benchmark, speed figure and quote traced to the lightonai blog or model cards. All performance and speed figures are attributed to 'the authors' and the article states that no independent reproduction exists in the sources. The blog date (October 8, 2026) is used, not the model-card update date (about Oct 9, shown as 'Updated 2 days ago'). Title claims verified: three sizes named in the sources are exactly 0.8B, 1B and 4B (the 1B appears in the blog, in the README and as LightOnOCR-3-1B on Hugging Face; it is not one of the two cited model cards, but the blog and README both document it). The title's size labels are the authors' names, not measured sizes: the 4B card shows 5B params (4.54 billion) and the 0.8B card shows 0.9B; the article discloses this under What We Don't Know. Apache 2.0 confirmed on both cited model cards and in the blog. 'Bounding Boxes, Image Descriptions and Chart Tables' matches the blog's grounding-mode description. Mistral OCR 4.1 omission: the blog tables include a MistralOCR4.1 row (olmOCR 81.9, ParseBench 68.2, fr-bench 54.1) which scores below all three LightOnOCR-3 variants on every benchmark. The article names only Infinity Parser Pro and Chandra 2 as comparators, never claims an exhaustive ranking, and states where Infinity Parser Pro and other models lead; the omission does not make the comparison one-sided (it would only strengthen the vendor's position), though ParseBench 'first and second' is the authors' claim over a ten-model table that the article does not enumerate. Judged acceptable. Quote in Overview and the short 'grounding' prompt quote are verbatim.
Overall Assessment: Accurate, well-attributed vendor-release coverage. Prompt-injection flags are false positives from embedded chat-template text. APPROVE.