AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apar

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:32:20.942Z. This is not the publication date.