AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families.

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T04:31:57.454Z. This is not the publication date.