Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment-Reference-Cited by-同舟云学术

Relative Entropy of Correct Proximal Policy Optimization Algorithms with Modified Penalty Factor in Complex Environment

Published:2022-03-22 Issue:4 Volume:24 Page:440
ISSN:1099-4300
Container-title:Entropy
language:en
Short-container-title:Entropy

Author:

Chen Weimin,Wong Kelvin Kian Loong,Long Sifan^ORCID,Sun Zhili

Abstract

In the field of reinforcement learning, we propose a Correct Proximal Policy Optimization (CPPO) algorithm based on the modified penalty factor β and relative entropy in order to solve the robustness and stationarity of traditional algorithms. Firstly, In the process of reinforcement learning, this paper establishes a strategy evaluation mechanism through the policy distribution function. Secondly, the state space function is quantified by introducing entropy, whereby the approximation policy is used to approximate the real policy distribution, and the kernel function estimation and calculation of relative entropy is used to fit the reward function based on complex problem. Finally, through the comparative analysis on the classic test cases, we demonstrated that our proposed algorithm is effective, has a faster convergence speed and better performance than the traditional PPO algorithm, and the measure of the relative entropy can show the differences. In addition, it can more efficiently use the information of complex environment to learn policies. At the same time, not only can our paper explain the rationality of the policy distribution theory, the proposed framework can also balance between iteration steps, computational complexity and convergence speed, and we also introduced an effective measure of performance using the relative entropy concept.

Publisher

MDPI AG

Subject

General Physics and Astronomy

Link

https://www.mdpi.com/1099-4300/24/4/440/pdf

Reference28 articles.

1. Mastering the game of Go without human knowledge

2. Benchmarking Deep Reinforcement Learning for Continuous Control;Yan;Proc. Mach. Learn. Res.,2016

3. Robot gains Social Intelligence through Multimodal Deep Reinforcement Learning;Hussain;arXiv,2017

4. Human-level control through deep reinforcement learning

5. Deep reinforcement learning: An overview;Li;arXiv,2017

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Research on Gait Switching Method Based on Speed Requirement;Journal of Bionic Engineering;2024-09-04

2. Efficient Difficulty Level Balancing in Match-3 Puzzle Games: A Comparative Study of Proximal Policy Optimization and Soft Actor-Critic Algorithms;Electronics;2023-10-30

3. AutoInfo GAN: Toward a better image synthesis GAN framework for high-fidelity few-shot datasets via NAS and contrastive learning;Knowledge-Based Systems;2023-09

4. Sustainable Coupling Coordination and Influencing Factors of Sports Facilities Construction and Social Economy Development in China;Sustainability;2023-02-03