LightOn, the Paris AI company listed on Euronext Growth, released LightOnOCR-3 on 8 October, a new family of models that turn documents into text and structured data. They come in two sizes, 0.8 billion and 4 billion parameters, and are published under the Apache 2.0 licence, which allows commercial use, modification and redistribution, according to the company's announcement and technical article.
More than text recognition
Classic OCR reads the words on a page. LightOnOCR-3 also identifies the elements of a document, such as paragraphs, headings, tables and images, and returns their bounding-box coordinates, so applications can link any answer back to the place on the original page. It describes images in a short text, and for charts and scientific figures it returns an HTML table of data points. A simple change of prompt switches between plain transcription and this richer output.
The company pitches it at three uses: financial analysis, where reports mix text, tables and charts; administrative workflows, where it claims the largest gains on French-language documents and handwriting; and technical document search for RAG and enterprise search.
The benchmarks, as LightOn reports them
On OlmOCR-Bench, LightOn says the model ranks second, behind Infinity-Parser2-Pro, a model more than eight times its size, and ahead of Chandra-OCR-2 and Mistral OCR 4.1. It says it leads open-weight models on ParseBench and ranks first on FRBench-pdf2md, a test built on French documents. The models process documents twice as fast as LightOnOCR-2, and when run on an organisation's own servers the company puts the cost below one cent per thousand pages, depending on hardware and workload.
The previous version, LightOnOCR-2, sees 240,000 monthly downloads on Hugging Face. LightOn was founded in Paris in 2016 and describes itself as the first European AI company listed on Euronext Growth; it sells an enterprise AI platform to the finance, industrial, healthcare, defence and public sectors.



