TeleOCR: a 1.2B open model that tops document parsing, runs local
StarDoc's Apache-2.0 vision model turns scanned and photographed pages into HTML tables and LaTeX, beating MinerU 2.5 Pro and rivaling far larger models.

Copy markdown
A 1.2B model that reads documents like a big one
StarDoc-AI's TeleOCR is a 1.2B vision-language model (built on Qwen2.5-VL) that parses digital and camera-captured documents in one pass — text, layout, reading order, tables and formulas. It reports 96.87 on OmniDocBench v1.6, state-of-the-art for its size class.
It out-parses MinerU 2.5 Pro, and coverage says Gemini 3 Pro
On the tougher Dr.DocBench challenge TeleOCR scores 67.96 to MinerU 2.5 Pro's 62.26, and it ranked #1 on ICDAR2026's Sci-ImageMiner. Writeups frame it as edging Gemini 3 Pro on document parsing — from a model small enough to run on your own GPU.
Tables to HTML, formulas to LaTeX, no preprocessing
Output is structured: tables in OTSL (convert straight to HTML), math in LaTeX, plus reading order. Geometry-aware modeling reads warped phone photos without a deskew step — a fully local doc-to-markdown pipeline you could wire up this weekend.
Apache-2.0, and it fits on a modest GPU
Weights ship Apache-2.0 as BF16 safetensors, so commercial use is unambiguous. At 1.2B it needs only a few GB of VRAM, runs on Transformers, vLLM or SGLang, and GGUF builds are out for Ollama, LM Studio and llama.cpp.