AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, comparing three in-context example retrieval strategies: domain-random selection, dense semantic retrieval, and our proposed CharBM25, which retrieves examples by character n-gram BM25 similarity over OCR inputs to target shared error patterns with the test sentence. Across a 20,000-sentence benchmark

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T13:51:27.104Z. This is not the publication date.