SOURCE-LINKED INTELLIGENCE
Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models
A language model with access to tools can commit to a final claim unsupported by the evidence it has seen, even when a single available tool call would resolve the uncertainty and its instructions explicitly forbid assumptions and guesses. We separate this failure into two precisely defined quantities: occurrence, how often the model makes an unsupported claim on its own, measured from the visible evidence and final claim without using the hidden correct answer; and conditional repair, how often those same naturally occurring unsupported claims are repaired when the missing evidence is supplie
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-27T23:02:52.000Z
First collected: 2026-09-21T08:21:55.975Z. This is not the publication date.