AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

List Counting Failures Are Not One Phenomenon

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Counting the items in a bracketed list looks trivial, yet open-weight chat models often get it wrong. Prior work usually blames input bottlenecks such as subword fragmentation or attention dilution, which predict that different models should fail in roughly the same way. Across seven instruct models on identical prompts, however, wrong answers form distinct modes: Qwen and Gemma 27B often flip odd lengths to a nearby even integer, OLMo concentrates errors on a few mid-sized integers, and Llama tends to under-count. These modes are useful labels rather than a stable family law (Gemma 9B does no

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T01:22:24.568Z. This is not the publication date.