AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Two Global Crops Suffice: Locating Semantic Emergence in DINO-Style Self-Supervised Learning

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

Self-supervised vision transformers trained with DINO-style objectives exhibit striking emergent semantic representation quality across visual tasks, yet the mechanisms underlying this behavior remain unclear. We present a systematic empirical dissection of the DINO family and show that semantic representations arise primarily from enforcing consistency between geometrically distinct global views of the same image instance. This instance-specific global alignment acts as the semantic anchor of DINO-style learning. Across controlled retraining experiments evaluated on semantic correspondence an

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T08:22:30.429Z. This is not the publication date.