Indoor Emergency Path Planning Based on the Q-Learning Optimization Algorithm-Reference-Cited by-同舟云学术

Indoor Emergency Path Planning Based on the Q-Learning Optimization Algorithm

Published:2022-01-14 Issue:1 Volume:11 Page:66
ISSN:2220-9964
Container-title:ISPRS International Journal of Geo-Information
language:en
Short-container-title:IJGI

Author:

Xu Shenghua,Gu Yang,Li Xiaoyan,Chen Cai,Hu Yingyi,Sang Yu,Jiang Wenxing

Abstract

The internal structure of buildings is becoming increasingly complex. Providing a scientific and reasonable evacuation route for trapped persons in a complex indoor environment is important for reducing casualties and property losses. In emergency and disaster relief environments, indoor path planning has great uncertainty and higher safety requirements. Q-learning is a value-based reinforcement learning algorithm that can complete path planning tasks through autonomous learning without establishing mathematical models and environmental maps. Therefore, we propose an indoor emergency path planning method based on the Q-learning optimization algorithm. First, a grid environment model is established. The discount rate of the exploration factor is used to optimize the Q-learning algorithm, and the exploration factor in the ε-greedy strategy is dynamically adjusted before selecting random actions to accelerate the convergence of the Q-learning algorithm in a large-scale grid environment. An indoor emergency path planning experiment based on the Q-learning optimization algorithm was carried out using simulated data and real indoor environment data. The proposed Q-learning optimization algorithm basically converges after 500 iterative learning rounds, which is nearly 2000 rounds higher than the convergence rate of the Q-learning algorithm. The SASRA algorithm has no obvious convergence trend in 5000 iterations of learning. The results show that the proposed Q-learning optimization algorithm is superior to the SARSA algorithm and the classic Q-learning algorithm in terms of solving time and convergence speed when planning the shortest path in a grid environment. The convergence speed of the proposed Q- learning optimization algorithm is approximately five times faster than that of the classic Q- learning algorithm. The proposed Q-learning optimization algorithm in the grid environment can successfully plan the shortest path to avoid obstacle areas in a short time.

Funder

National Key Research and Development Plan of China

Publisher

MDPI AG

Subject

Earth and Planetary Sciences (miscellaneous),Computers in Earth Sciences,Geography, Planning and Development

Link

https://www.mdpi.com/2220-9964/11/1/66/pdf

Reference40 articles.

1. Severity of the fire status and the impendency of establishing the related courses;Wu;Fire Sci. Technol.,2005

2. Security question in the fire evacuation;Meng;China Public Secur.,2005

3. 3D building information model for facilitating dynamic analysis of indoor fire emergency;Zhu;Geomat. Inf. Sci. Wuhan Univ.,2014

Cited by 11 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Artificial intelligence methodologies for building evacuation plan modeling;Journal of Building Engineering;2024-11

2. Cross-regional path planning based on improved Q-learning with dynamic exploration factor and heuristic reward value;Expert Systems with Applications;2024-09

3. An ensemble of brain storm optimization and Q-learning methods for distributed flexible job shop scheduling problems with distribution operations;International Journal of General Systems;2024-03-15

4. The Path Planning Research for Mobile Robot Based on Reinforcement Learning Particle Swarm Algorithm;2024 7th International Conference on Advanced Algorithms and Control Engineering (ICAACE);2024-03-01

5. The Wide-Area Coverage Path Planning Strategy for Deep-Sea Mining Vehicle Cluster Based on Deep Reinforcement Learning;Journal of Marine Science and Engineering;2024-02-12