AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VISTA: Dense Multi-Label Classroom Coding with Vision-Language Models

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Video-language benchmarks are usually constructed by the dataset authors without published reliability statistics, leaving the noise floor of the construct unknown. We argue that multimodal benchmarking benefits from methods taken from research communities that have already invested in strategies to ensure reliability. We illustrate the case with the Classroom Observation Protocol for Undergraduate STEM (COPUS): a 24-code multi-label observation instrument with a decade of peer-reviewed reliability literature. We recast COPUS as a video benchmark for multimodal foundation models, where it prov

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T04:31:57.454Z. This is not the publication date.