TeleOCR: a 1.2B open model that tops document parsing, runs local

StarDoc's Apache-2.0 vision model turns scanned and photographed pages into HTML tables and LaTeX, beating MinerU 2.5 Pro and rivaling far larger models.

Nowline OCT 1 3:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • A 1.2B model that reads documents like a big one

    StarDoc-AI's TeleOCR is a 1.2B vision-language model (built on Qwen2.5-VL) that parses digital and camera-captured documents in one pass — text, layout, reading order, tables and formulas. It reports 96.87 on OmniDocBench v1.6, state-of-the-art for its size class.

  • It out-parses MinerU 2.5 Pro, and coverage says Gemini 3 Pro

    On the tougher Dr.DocBench challenge TeleOCR scores 67.96 to MinerU 2.5 Pro's 62.26, and it ranked #1 on ICDAR2026's Sci-ImageMiner. Writeups frame it as edging Gemini 3 Pro on document parsing — from a model small enough to run on your own GPU.

  • Tables to HTML, formulas to LaTeX, no preprocessing

    Output is structured: tables in OTSL (convert straight to HTML), math in LaTeX, plus reading order. Geometry-aware modeling reads warped phone photos without a deskew step — a fully local doc-to-markdown pipeline you could wire up this weekend.

  • Apache-2.0, and it fits on a modest GPU

    Weights ship Apache-2.0 as BF16 safetensors, so commercial use is unambiguous. At 1.2B it needs only a few GB of VRAM, runs on Transformers, vLLM or SGLang, and GGUF builds are out for Ollama, LM Studio and llama.cpp.