SOURCE-LINKED INTELLIGENCE
ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines
We present ChunkRank, an open-source Python library that derives chunk boundaries from a target model's tokenizer and context window, and selects an answer among candidates produced independently per chunk. It ships a validated registry of 90 models across 15 providers and six answer-selection methods, and needs only three core dependencies. For chunking, ChunkRank avoids context-window overflow automatically from the model name, whereas character-based splitters overflow or waste the budget, and a fidelity study across 11 languages shows why token-exact budgets matter beyond English. For answ
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-24T13:59:51.000Z
First collected: 2026-09-25T06:12:46.948Z. This is not the publication date.