TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6
Original titleDigital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. 📜 Apache 2.0 License.
AISummary
TeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model.
It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge.
It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.
Source: ModelScope · x.comPublished · added here