AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverge from planner behavior, while planning distributions stored in replay become stale as the model and value function evolve. Reanalysis can refresh these targets, but at substantial computational cost. We propose GEM-MPC, an MPPI-based reinforcement learning method that improves t

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T13:51:27.104Z. This is not the publication date.