AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

arXiv · AI, language, vision and robotics · article · Oct 7, 2026 · UTC

Linking people's appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes. A strong vision-language model generates identity-aware captions and questions, which are manua

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-10-08T10:02:25.305Z. This is not the publication date.