AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduc

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T06:11:12.848Z. This is not the publication date.