AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Attributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR Comparison

arXiv · AI, language, vision and robotics · article · Sep 19, 2026 · UTC

A RAG system can hand you the right answer and cite a source it did not actually use. Models output these unfaithful citations via post-rationalization: they write the answer first and then attach a citation to whatever passage looks close enough. Search agents are now trained with reinforcement learning from verifiable rewards (RLVR), which pays them for getting the answer right. We asked whether that training also teaches them to cite honestly. Improving an existing methodology with a required control, we compared an instruction-tuned model against three RLVR agents trained from it, on four

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T10:01:48.231Z. This is not the publication date.