A novel multi-scale violence and public gathering dataset for crowd behavior classification-Reference-Cited by-同舟云学术

A novel multi-scale violence and public gathering dataset for crowd behavior classification

Published:2024-05-10 Issue: Volume:6 Page:
ISSN:2624-9898
Container-title:Frontiers in Computer Science
language:
Short-container-title:Front. Comput. Sci.

Author:

Elzein Almiqdad,Basaran Emrah,Yang Yin David,Qaraqe Marwa

Abstract

Dependable utilization of computer vision applications, such as smart surveillance, requires training deep learning networks on datasets that sufficiently represent the classes of interest. However, the bottleneck in many computer vision applications lies in the limited availability of adequate datasets. One particular application that is of great importance for the safety of cities and crowded areas is smart surveillance. Conventional surveillance methods are reactive and often ineffective in enable real-time action. However, smart surveillance is a key component of smart and proactive security in a smart city. Motivated by a smart city application which aims at the automatic identification of concerning events for alerting law-enforcement and governmental agencies, we craft a large video dataset that focuses on the distinction between small-scale violence, large-scale violence, peaceful gatherings, and natural events. This dataset classifies public events along two axes, the size of the crowd observed and the level of perceived violence in the crowd. We name this newly-built dataset the Multi-Scale Violence and Public Gathering (MSV-PG) dataset. The videos in the dataset go through several pre-processing steps to prepare them to be fed into a deep learning architecture. We conduct several experiments on the MSV-PG datasets using a ResNet3D, a Swin Transformer and an R(2 + 1)D architecture. The results achieved by these models when trained on the MSV-PG dataset, 88.37%, 89.76%, and 89.3%, respectively, indicate that the dataset is well-labeled and is rich enough to train deep learning models for automatic smart surveillance for diverse scenarios.

Funder

Qatar National Research Fund

Publisher

Frontiers Media SA

Reference69 articles.

1. “Vision-based fight detection from surveillance cameras,”;Aktı,2019

2. 3d-cnn-based fused feature maps with LSTM applied to action recognition;Arif;Future Internet,2019

3. “Vivit: a video vision transformer,”;Arnab,2021

4. “Violence detection in video using computer vision techniques,”;Bermejo Nievas;International Conference on Computer Analysis of Images and Patterns,2011