Affiliation:
1. School of Computer Engineering, Jimei University, Xiamen 361021, China
2. Xiamen Kingtop Information Technology Co., Ltd., Xiamen 361008, China
Abstract
Achieving precise individual localization within densely crowded scenes poses a significant challenge due to the intricate interplay of occlusions and varying density patterns. Traditional methods for crowd localization often rely on convolutional neural networks (CNNs) to generate density maps. However, these approaches are prone to inaccuracies stemming from the extensive overlaps inherent in dense populations. To overcome this challenge, our study introduces the Hierarchical Inverse Distance Transformer (HIDT), a novel framework that harnesses the multi-scale global receptive fields of Pyramid Vision Transformers. By adapting to the multi-scale characteristics of crowds, HIDT significantly enhances the accuracy of individual localization. Incorporating Focal Inverse Distance techniques, HIDT adeptly addresses issues related to scale variation and dense overlaps, prioritizing local small-scale features within the broader contextual understanding of the scene. Rigorous evaluation on standardized benchmarks has unequivocally validated the superiority of our approach. HIDT exhibits outstanding performance across various datasets. Notably, on the JHU-Crowd++ dataset, our method demonstrates significant improvements over the baseline, with MAE and MSE metrics decreasing from 66.6 and 253.6 to 59.1 and 243.5, respectively. Similarly, on the UCF-QNRF dataset, performance metrics increase from 89.0 and 153.5 to 83.6 and 138.7, highlighting the effectiveness and versatility of our approach.
Funder
Natural Science Foundation of Xiamen, China
National Natural Science Foundation of China
Natural Science Foundation of Fujian Province
Reference28 articles.
1. Abousamra, S., Hoai, M., Samaras, D., and Chen, C. (2021, January 2–9). Localization in the crowd with topological constraints. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually.
2. Liu, Y., Shi, M., Zhao, Q., and Wang, X. (2019, January 16–17). Point in, box out: Beyond counting persons in crowds. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.
3. Locate, size, and count: Accurately resolving people in dense crowds via detection;Sam;IEEE Trans. Pattern Anal. Mach. Intell.,2020
4. Faster R-CNN: Towards real-time object detection with region proposal networks;Ren;IEEE Trans. Pattern Anal. Mach. Intell.,2016
5. Focal inverse distance transform maps for crowd localization;Liang;IEEE Trans. Multimed.,2022