AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Understanding Reliability in LLM-based Human Behavior Simulation

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completene

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T18:11:26.115Z. This is not the publication date.