AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

CantoneseLLM v2: Reasoning in a Low-Resource Language

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

Cantonese is widely spoken but remains low-resource in written data, with no large corpus of native Cantonese reasoning traces available for model training. We develop and release CantoneseLLM v2, comprising models based on Qwen3 8B and 30B-A3B. The models are trained through CPT on 784 million Cantonese and Hong Kong-related tokens, chat-vector merging, SFT, DPO, and RLVR. Evaluation across the training stages shows that chat-vector merging transfers instruction following but preserves the donor model's reasoning language, while SFT with limited Cantonese reasoning data substantially shortens

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:52:10.320Z. This is not the publication date.