AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

arXiv · AI, language, vision and robotics · article · Sep 17, 2026 · UTC

We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so, a Limit-Deterministic B{ü}chi Automaton (LDBA) representation of the LTL task is synchronised with a Bayes-Adaptive Markov Decision Process (BAMDP) representation of the environment, which allows us to leverage an enhanced exploration-exploitation trade-off that is achieved via Bayesian RL, as opposed to traditional non-Bayesian approaches. We further propose a nove

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:12:08.350Z. This is not the publication date.