AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

How Do Language Models Represent and Use Phonological Information for Allomorph Selection?

arXiv · AI, language, vision and robotics · article · Sep 4, 2026 · UTC

Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably produce morphemes whose form is phonologically conditioned. It remains unclear whether they rely on item-specific memorization or rule-like generalization and, if the latter, how that generalization is implemented. We therefore ask whether this phonological condition is represented within language models and how it is causally used for allomorph selection. For the English indefinite article a/an, we show that the phonological condition is encoded along a single linear direction in trigge

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T22:31:48.298Z. This is not the publication date.