SOURCE-LINKED INTELLIGENCE
Double Trouble: Bilingual Pretraining Leaves Language-Conditioned Effects in Shared-Language Representations
A concept can carry different associations across languages, while modern language models learn English alongside many other languages during pretraining. Yet comparisons among existing models cannot easily isolate how any one language changes the way these models represent English concepts because their training corpora, compute, architectures, and random seeds all differ. We study this question through a controlled experiment with 40 matched 310M-parameter decoder-only models that share an architecture, tokenizer, training recipe, and English data source. Each bilingual condition adds one of
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-27T03:39:24.000Z
First collected: 2026-09-21T08:51:59.673Z. This is not the publication date.