An Improved Dueling Deep Double-Q Network Based on Prioritized Experience Replay for Path Planning of Unmanned Surface Vehicles-Reference-Cited by-同舟云学术

An Improved Dueling Deep Double-Q Network Based on Prioritized Experience Replay for Path Planning of Unmanned Surface Vehicles

Published:2021-11-13 Issue:11 Volume:9 Page:1267
ISSN:2077-1312
Container-title:Journal of Marine Science and Engineering
language:en
Short-container-title:JMSE

Author:

Zhu Zhengwei,Hu Can,Zhu Chenyang^ORCID,Zhu Yanping,Sheng Yu

Abstract

Unmanned Surface Vehicle (USV) has a broad application prospect and autonomous path planning as its crucial technology has developed into a hot research direction in the field of USV research. This paper proposes an Improved Dueling Deep Double-Q Network Based on Prioritized Experience Replay (IPD3QN) to address the slow and unstable convergence of traditional Deep Q Network (DQN) algorithms in autonomous path planning of USV. Firstly, we use the deep double Q-Network to decouple the selection and calculation of the target Q value action to eliminate overestimation. The prioritized experience replay method is adopted to extract experience samples from the experience replay unit, increase the utilization rate of actual samples, and accelerate the training speed of the neural network. Then, the neural network is optimized by introducing a dueling network structure. Finally, the soft update method is used to improve the stability of the algorithm, and the dynamic ϵ-greedy method is used to find the optimal strategy. The experiments are first conducted in the Open AI Gym test platform to pre-validate the algorithm for two classical control problems: the Cart pole and Mountain Car problems. The impact of algorithm hyperparameters on the model performance is analyzed in detail. The algorithm is then validated in the Maze environment. The comparative analysis of simulation experiments shows that IPD3QN has a significant improvement in learning performance regarding convergence speed and convergence stability compared with DQN, D3QN, PD2QN, PDQN, PD3QN. Also, USV can plan the optimal path according to the actual navigation environment with the IPD3QN algorithm.

Funder

Changzhou Key Research and Development Program

Publisher

MDPI AG

Subject

Ocean Engineering,Water Science and Technology,Civil and Structural Engineering

Link

https://www.mdpi.com/2077-1312/9/11/1267/pdf

Reference26 articles.

1. Applications of marine robotic vehicles

2. Multi-Robot Path Planning Method Using Reinforcement Learning

3. A novel reinforcement learning based grey wolf optimizer algorithm for unmanned aerial vehicles (UAVs) path planning

4. Human-level control through deep reinforcement learning

Cited by 16 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Optimization of 2D Irregular Packing: Deep Reinforcement Learning with Dense Reward;International Journal of Semantic Computing;2024-05-27

2. Deep reinforcement learning algorithm based ramp merging decision model;Proceedings of the Institution of Mechanical Engineers, Part D: Journal of Automobile Engineering;2024-03-29

3. ChemGymRL: A customizable interactive framework for reinforcement learning for digital chemistry;Digital Discovery;2024

4. Path planning of USV based on improved PRM under the influence of ocean current;Proceedings of the Institution of Mechanical Engineers, Part M: Journal of Engineering for the Maritime Environment;2023-12-15

5. An AUV-Assisted Data Gathering Scheme Based on Deep Reinforcement Learning for IoUT;Journal of Marine Science and Engineering;2023-11-30