01407nas a2200205 4500000000100000000000100001008004100002260002100043100001900064700002400083700002400107700002000131700002000151700002300171700002000194700001800214245007500232856004600307520084800353 2023 d aMontreal, Canada1 aSamuel Kessler1 aMateusz Ostaszewski1 aMichał Bortkiewicz1 aMateusz Żarski1 aMaciej Wołczyk1 aJack Parker-Holder1 aStephen Roberts1 aPiotr Miłoś00aThe Effectiveness of World Models for Continual Reinforcement Learning uhttps://lifelong-ml.cc/online_proceedings3 a
World models power some of the most efficient reinforcement learning algorithms. In this work, we showcase that they can be harnessed for continual learning – a situation when the agent faces changing environments. World models typically employ a replay buffer for training, which can be naturally extended to continual learning. We systematically study how different selective experience replay methods affect performance, forgetting, and transfer. We also provide recommendations regarding various modeling options for using world models. The best set of choices is called Continual-Dreamer, it is task-agnostic and utilizes the world model for continual exploration. Continual-Dreamer is sample efficient and outperforms state-of-the-art task-agnostic continual reinforcement learning methods on Minigrid and Minihack benchmarks.