AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

arXiv · AI, language, vision and robotics · article · Aug 27, 2026 · UTC

Zero-shot classification needs efficient label retrieval and fine-grained visual reasoning, yet discriminative and generative vision-language models fail in complementary ways.When CLIP's top-1 prediction is wrong, the correct label often remains in its top-$K$ shortlist, making disambiguation rather than recall the key challenge.Standalone generative models, however, are hindered by large label spaces and unconstrained outputs.This complementarity motivates separating broad candidate retrieval from fine-grained, image-grounded verification.We propose G2D, a training-free framework that uses a

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:51:59.673Z. This is not the publication date.