AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin

arXiv · AI, language, vision and robotics · article · Sep 5, 2026 · UTC

Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The best labelled attachment score is 0.62 and the best morphology-aware score is 0.24. Performance does not correlate with either genre or period proximity. To address this shortfall, in-domain training data was generated as a by-product of using these inadequate models. In each of nine iterations, a model pre-annotated 200 senten

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:32:07.623Z. This is not the publication date.