Time-aware MADDPG with LSTM for multi-agent obstacle avoidance: a comparative study-Reference-Cited by-同舟云学术

Time-aware MADDPG with LSTM for multi-agent obstacle avoidance: a comparative study

Published:2024-03-02 Issue:3 Volume:10 Page:4141-4155
ISSN:2199-4536
Container-title:Complex & Intelligent Systems
language:en
Short-container-title:Complex Intell. Syst.

Author:

Zhao Enyu,Zhou Ning,Liu Chanjuan^ORCID,Su Houfu,Liu Yang,Cong Jinmiao

Abstract

AbstractIntelligent agents and multi-agent systems are increasingly used in complex scenarios, such as controlling groups of drones and non-player characters in video games. In these applications, multi-agent navigation and obstacle avoidance are foundational functions. However, problems become more challenging with the increased complexity of the environment and the dynamic decision-making interactions among agents. The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is a classical multi-agent reinforcement learning algorithm successfully used to improve agents’ performance. However, it ignores the temporal message hidden in agents’ interaction with the environment and needs to be more efficient in scenarios with many agents due to its training technique. To address the limitations of MADDPG, we propose to explore modified algorithms of MADDPG for multi-agent navigation and obstacle avoidance. By combining MADDPG with Long Short-Term Memory (LSTM), we obtain the MADDPG-LSTMactor algorithm, which leverages continuous observations over time as input for the policy network, enabling the LSTM layer to capture hidden temporal patterns. Moreover, by simplifying the input of the critic network, we obtain the MADDPG-L algorithm for efficiency improvement in scenarios with many agents. Experimental results demonstrate that these algorithms outperform existing networks in the OpenAI multi-agent particle environment. We also conducted a comparative study of the LSTM-based approach with Transformer and self-attention models in the task of multi-agent navigation and obstacle avoidance. The results reveal that Transformer and self-attention do not consistently outperform LSTM. The LSTM-based model exhibits a favorable tradeoff across varying sequence lengths. Overall, this work addresses the limitations of MADDPG in multi-agent navigation and obstacle avoidance tasks, providing insights for developing intelligent agents and multi-agent systems.

Funder

National Natural Science Foundation of China

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s40747-024-01389-0.pdf

Reference20 articles.

1. Zhong J, Wang T, Cheng L (2022) Collision-free path planning for welding manipulator via hybrid algorithm of deep reinforcement learning and inverse kinematics. Complex Intell Syst 8:1899–1912

2. Dechter R, Pearl J (1985) Generalized best-first search strategies and the optimality of A. J ACM 32(3):505–536

3. Tan M (1993) Multi-agent reinforcement learning: independent vs. cooperative agents. Machine Learning Proceedings, pp 330–337

4. Watkins CJCH, Dayan P (1992) Q-learning. Mach Learn 8(3):279–292

5. de Witt CS, Gupta T, Makoviichuk D, et al. (2020) Is independent learning all you need in the starcraft multi-agent challenge? arXiv preprint arXiv:2011.09533. Accessed 15 Dec 2023.

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Research on Cooperative Obstacle Avoidance Decision Making of Unmanned Aerial Vehicle Swarms in Complex Environments under End-Edge-Cloud Collaboration Model;Drones;2024-09-04

2. RNN-LSTM: From applications to modeling techniques and beyond—Systematic review;Journal of King Saud University - Computer and Information Sciences;2024-06

3. Boosting Reinforcement Learning via Hierarchical Game Playing With State Relay;IEEE Transactions on Neural Networks and Learning Systems;2024