AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation, OCR outputs are strongly grounded in the input image, making diffusion-based parallel generation promising. However, when several tokens are predicted in one diffusion step, each is predicted before the others are known. Committing them directly can therefore introduce errors. We therefore introduce GravityOCR, a parameter-shared AR-block-diffusion model jointly train

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:11:12.117Z. This is not the publication date.