AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this work, we propose a decoupled forecasting framework that separates future representation learning from video generation. A frozen V-JEPA encoder first extracts predictive latent representations from the observed traffic videos, capturing the underlying scene dynamics in a semantic la

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:01:59.594Z. This is not the publication date.