Pyyan / OCR & Document AI

OCR & Document AI

12 entries

Turning pages into structured text. Now dominated by vision language models rather than the classical OCR engines. · ranked by document parsing benchmarks · verified 13 Aug 2026

#NameOmniDocBenchOpen weightsHandlesLicenceKindStatus
1GLM-OCR
Zhipu AI
94.6YesTables, formulas, handwritingOpen weightsVision language modelcurrent
2dots.ocr
Xiaohongshu
~93YesLayout, 100+ languagesMITVision language modelcurrent
3Qwen3-VL
Alibaba
~93YesDocuments, charts, videoApache 2.0Vision language modelcurrent
4DeepSeek-OCR
DeepSeek
~92YesDense text, tablesMITVision language modelcurrent
5Mistral OCR
Mistral AI
~92NoTables, images, equationsProprietaryHosted APIcurrent
6olmOCR
Allen Institute for AI
~91YesTables, markdown structureApache 2.0Vision language modelcurrent
7GOT-OCR 2.0
StepFun
~90YesFormulas, music, chartsApache 2.0Vision language modelcurrent
8MinerU
OpenDataLab
~90YesPDF, formulas, tablesAGPL-3.0Pipelinecurrent
9Docling
IBM
~88YesPDF, DOCX, PPTX, HTMLMITPipelinecurrent
10Surya
Datalab
~87YesLayout, reading order, tablesGPL-3.0Pipelinecurrent
11PaddleOCR
Baidu
~82YesText, tables, receiptsApache 2.0Classical enginecurrent
12Tesseract
Google
~55YesClean printed textApache 2.0Classical enginecurrent