AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Auditing Harness Tampering in Self-Improving Agents

arXiv · AI, language, vision and robotics · article · Aug 30, 2026 · UTC

Self-improving agents iteratively modify their own harness to push the frontier of their performance. However, such modifications can produce illusory performance gains or compromise integrity constraints such as authorization, provenance, and completeness without genuinely improving capability. We term this phenomenon as harness tampering, which extends the concept from reward and measurement tampering to the full self-improvement lifecycle. To systematically study this problem, we propose a two-axis taxonomy that categorizes each misaligned edit by the harness functional role in which it occ

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:22:03.933Z. This is not the publication date.