SOURCE-LINKED INTELLIGENCE
InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. It lifts patch features into a continuous spatial representation, propagates directed influence over multiple steps, and predicts local intervention effects through a shared transition operator. Training jointly optimizes language modeling, cross-environment invariance, counterfactual rollout supervisio
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-07T18:36:20.000Z
First collected: 2026-09-20T20:32:20.942Z. This is not the publication date.