AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which require reasoning over semantic perturbations of objects, attributes, relations, and their interactions. However, our controlled analysis reveals that existing compositionality-aware VLMs exhibit element-specific biases, often underperforming vanilla CLIP on certain compositional elements. To address this, we propose Compositional Scene Graph-guided CLIP (CS-CLIP), which uses scene graphs to identify compositional elements and construct structured negatives via selective masking. We fu

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:22:01.598Z. This is not the publication date.