AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environm

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.