{"slug": "jina-ocr-v1-efficient-document-parsing-with-speculative-decoding-and-dense", "title": "Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards", "summary": "Jina AI released Jina-OCR-v1, an end-to-end document parsing model combining DeepSeek-OCR's compressed-vision encoder and 3B mixture-of-experts decoder (activating ~570M parameters per token) with a FastMTP speculative decoding head. The model scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, achieving 2.57 pages per second throughput, and doubles decoding speed on NVIDIA L4 GPUs. It is publicly available on Hugging Face.", "body_md": "arXiv:2609.03181v1 Announce Type: new\nAbstract: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at https://huggingface.co/jinaai/jina-ocr-v1.", "url": "https://wpnews.pro/news/jina-ocr-v1-efficient-document-parsing-with-speculative-decoding-and-dense", "canonical_source": "https://arxiv.org/abs/2609.03181", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:22:40.747767+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-products"], "entities": ["Jina AI", "DeepSeek-OCR", "FastMTP", "NVIDIA L4", "OmniDocBench v1.6", "olmOCR-Bench", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/jina-ocr-v1-efficient-document-parsing-with-speculative-decoding-and-dense", "markdown": "https://wpnews.pro/news/jina-ocr-v1-efficient-document-parsing-with-speculative-decoding-and-dense.md", "text": "https://wpnews.pro/news/jina-ocr-v1-efficient-document-parsing-with-speculative-decoding-and-dense.txt", "jsonld": "https://wpnews.pro/news/jina-ocr-v1-efficient-document-parsing-with-speculative-decoding-and-dense.jsonld"}}