SOURCE-LINKED INTELLIGENCE
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environm
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-18T18:42:03.000Z
First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.