Transformers for Urban Sound Classification—A Comprehensive Performance Evaluation-Reference-Cited by-同舟云学术

Transformers for Urban Sound Classification—A Comprehensive Performance Evaluation

Published:2022-11-16 Issue:22 Volume:22 Page:8874
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Nogueira Ana Filipa Rodrigues^ORCID,Oliveira Hugo S.^ORCID,Machado José J. M.^ORCID,Tavares João Manuel R. S.^ORCID

Abstract

Many relevant sound events occur in urban scenarios, and robust classification models are required to identify abnormal and relevant events correctly. These models need to identify such events within valuable time, being effective and prompt. It is also essential to determine for how much time these events prevail. This article presents an extensive analysis developed to identify the best-performing model to successfully classify a broad set of sound events occurring in urban scenarios. Analysis and modelling of Transformer models were performed using available public datasets with different sets of sound classes. The Transformer models’ performance was compared to the one achieved by the baseline model and end-to-end convolutional models. Furthermore, the benefits of using pre-training from image and sound domains and data augmentation techniques were identified. Additionally, complementary methods that have been used to improve the models’ performance and good practices to obtain robust sound classification models were investigated. After an extensive evaluation, it was found that the most promising results were obtained by employing a Transformer model using a novel Adam optimizer with weight decay and transfer learning from the audio domain by reusing the weights from AudioSet, which led to an accuracy score of 89.8% for the UrbanSound8K dataset, 95.8% for the ESC-50 dataset, and 99% for the ESC-10 dataset, respectively.

Funder

the project Safe Cities—”Inovação para Construir Cidades Seguras”

the European Regional Development Fund

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/22/22/8874/pdf

Reference30 articles.

1. Virtanen, T., Plumbley, M.D., and Ellis, D. (2018). Computational Analysis of Sound Scenes and Events, Springer International Publishing.

2. Zinemanas, P., Rocamora, M., Miron, M., Font, F., and Serra, X. (2021). An Interpretable Deep Learning Model for Automatic Sound Classification. Electronics, 10.

3. Environmental sound classification using convolution neural networks with different integrated loss functions;Expert Syst.,2021

4. Das, J.K., Ghosh, A., Pal, A.K., Dutta, S., and Chakrabarty, A. (2020, January 21–23). Urban Sound Classification Using Convolutional Neural Network and Long Short Term Memory Based on Multiple Features. Proceedings of the 2020 Fourth International Conference on Intelligent Computing in Data Sciences (ICDS), Fez, Morocco.

5. Mushtaq, Z., and Su, S.F. (2020). Efficient Classification of Environmental Sounds through Multiple Features Aggregation and Data Enhancement Techniques for Spectrogram Images. Symmetry, 12.

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Graph-Based Audio Classification Using Pre-Trained Models and Graph Neural Networks;Sensors;2024-03-26

2. Invisible track bed defect classification method based on distributed optical fiber sensing system and FFT Attention Transformer model;Optical Engineering;2024-02-20

3. Cross-modal and Cross-medium Adversarial Attack for Audio;Proceedings of the 31st ACM International Conference on Multimedia;2023-10-26

4. Misophonia Sound Recognition Using Vision Transformer;2023 45th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC);2023-07-24

5. A Comparative Analysis of Image Captioning Techniques;2023 4th International Conference for Emerging Technology (INCET);2023-05-26