On Entropy Regularized Path Integral Control for Trajectory Optimization-Reference-Cited by-同舟云学术

On Entropy Regularized Path Integral Control for Trajectory Optimization

Published:2020-10-03 Issue:10 Volume:22 Page:1120
ISSN:1099-4300
Container-title:Entropy
language:en
Short-container-title:Entropy

Author:

Lefebvre Tom^ORCID,Crevecoeur Guillaume

Abstract

In this article, we present a generalized view on Path Integral Control (PIC) methods. PIC refers to a particular class of policy search methods that are closely tied to the setting of Linearly Solvable Optimal Control (LSOC), a restricted subclass of nonlinear Stochastic Optimal Control (SOC) problems. This class is unique in the sense that it can be solved explicitly yielding a formal optimal state trajectory distribution. In this contribution, we first review the PIC theory and discuss related algorithms tailored to policy search in general. We are able to identify a generic design strategy that relies on the existence of an optimal state trajectory distribution and finds a parametric policy by minimizing the cross-entropy between the optimal and a state trajectory distribution parametrized by a parametric stochastic policy. Inspired by this observation, we then aim to formulate a SOC problem that shares traits with the LSOC setting yet that covers a less restrictive class of problem formulations. We refer to this SOC problem as Entropy Regularized Trajectory Optimization. The problem is closely related to the Entropy Regularized Stochastic Optimal Control setting which is often addressed lately by the Reinforcement Learning (RL) community. We analyze the theoretical convergence behavior of the theoretical state trajectory distribution sequence and draw connections with stochastic search methods tailored to classic optimization problems. Finally we derive explicit updates and compare the implied Entropy Regularized PIC with earlier work in the context of both PIC and RL for derivative-free trajectory optimization.

Funder

vlaamse overheid

Publisher

MDPI AG

Subject

General Physics and Astronomy

Link

https://www.mdpi.com/1099-4300/22/10/1120/pdf

Reference65 articles.

1. Emergence of locomotion behaviours in rich environments;Heess;arXiv,2017

2. Optimal control theory;Todorov,2006

3. A Second-order Gradient Method for Determining Optimal Trajectories of Non-linear Discrete-time Systems

4. Control-limited differential dynamic programming

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Probabilistic control and majorisation of optimal control;Systems & Control Letters;2024-08

2. Dual regularized policy updating and shiftpoint detection for automated deployment of reinforcement learning controllers on industrial mechatronic systems;Control Engineering Practice;2024-01

3. Entropy Regularised Deterministic Optimal Control: From Path Integral Solution to Sample-Based Trajectory Optimisation;2022 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM);2022-07-11

4. Exploratory LQG mean field games with entropy regularization;Automatica;2022-05

5. Trajectory Planning of Robot Manipulator Based on RBF Neural Network;Entropy;2021-09-13