AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases

arXiv · AI, language, vision and robotics · article · Sep 24, 2026 · UTC

Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text that has to be coded before it can be counted, and the result has to be legible and rigorous enough to meaningfully compare models and vendors. To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior. Each study is a frozen, public stimulus run identically on a cross-vendor panel, at a few dollars per mode

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T06:12:46.948Z. This is not the publication date.