YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design-Reference-Cited by-同舟云学术

YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-Design

Published:2021-05-18 Issue:2 Volume:35 Page:955-963
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Cai Yuxuan,Li Hongjia,Yuan Geng,Niu Wei,Li Yanyu,Tang Xulong,Ren Bin,Wang Yanzhi

Abstract

The rapid development and wide utilization of object detection techniques have aroused attention on both accuracy and speed of object detectors. However, the current state-of-the-art object detection works are either accuracy-oriented using a large model but leading to high latency or speed-oriented using a lightweight model but sacrificing accuracy. In this work, we propose YOLObile framework, a real-time object detection on mobile devices via compression-compilation co-design. A novel block-punched pruning scheme is proposed for any kernel size. To improve computational efficiency on mobile devices, a GPU-CPU collaborative scheme is adopted along with advanced compiler-assisted optimizations. Experimental results indicate that our pruning scheme achieves 14x compression rate of YOLOv4 with 49.0 mAP. Under our YOLObile framework, we achieve 17 FPS inference speed using GPU on Samsung Galaxy S20. By incorporating our proposed GPU-CPU collaborative scheme, the inference speed is increased to 19.1 FPS, and outperforms the original YOLOv4 by 5x speedup. Source code is at: https://github.com/nightsnack/YOLObile.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 36 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ESFD-YOLOv8n: Early Smoke and Fire Detection Method Based on an Improved YOLOv8n Model;Fire;2024-08-27

2. A comprehensive survey of deep learning-based lightweight object detection models for edge devices;Artificial Intelligence Review;2024-08-10

3. FPGA-based CNN Acceleration using Pattern-Aware Pruning;2024 IEEE 6th International Conference on AI Circuits and Systems (AICAS);2024-04-22

4. Towards lightweight military object detection;Journal of Intelligent & Fuzzy Systems;2024-04-18

5. Scalable Solutions for Efficient Real-Time Distributed Video Analytics with Vehicle Detection on CPU Edge Nodes;Proceedings of the 2024 7th International Conference on Computers in Management and Business;2024-01-12