AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables,

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:32:20.942Z. This is not the publication date.