SOURCE-LINKED INTELLIGENCE
Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
Linking people's appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes. A strong vision-language model generates identity-aware captions and questions, which are manua
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-10-07T14:41:29.000Z
First collected: 2026-10-08T10:02:25.305Z. This is not the publication date.