Efficient Classification of Environmental Sounds through Multiple Features Aggregation and Data Enhancement Techniques for Spectrogram Images-Reference-Cited by-同舟云学术

Efficient Classification of Environmental Sounds through Multiple Features Aggregation and Data Enhancement Techniques for Spectrogram Images

Published:2020-11-03 Issue:11 Volume:12 Page:1822
ISSN:2073-8994
Container-title:Symmetry
language:en
Short-container-title:Symmetry

Author:

Mushtaq Zohaib^ORCID,Su Shun-Feng

Abstract

Over the past few years, the study of environmental sound classification (ESC) has become very popular due to the intricate nature of environmental sounds. This paper reports our study on employing various acoustic features aggregation and data enhancement approaches for the effective classification of environmental sounds. The proposed data augmentation techniques are mixtures of the reinforcement, aggregation, and combination of distinct acoustics features. These features are known as spectrogram image features (SIFs) and retrieved by different audio feature extraction techniques. All audio features used in this manuscript are categorized into two groups: one with general features and the other with Mel filter bank-based acoustic features. Two novel and innovative features based on the logarithmic scale of the Mel spectrogram (Mel), Log (Log-Mel) and Log (Log (Log-Mel)) denoted as L2M and L3M are introduced in this paper. In our study, three prevailing ESC benchmark datasets, ESC-10, ESC-50, and Urbansound8k (Us8k) are used. Most of the audio clips in these datasets are not fully acquired with sound and include silence parts. Therefore, silence trimming is implemented as one of the pre-processing techniques. The training is conducted by using the transfer learning model DenseNet-161, which is further fine-tuned with individual optimal learning rates based on the discriminative learning technique. The proposed methodologies attain state-of-the-art outcomes for all used ESC datasets, i.e., 99.22% for ESC-10, 98.52% for ESC-50, and 97.98% for Us8k. This work also considers real-time audio data to evaluate the performance and efficiency of the proposed techniques. The implemented approaches also have competitive results on real-time audio data.

Publisher

MDPI AG

Subject

Physics and Astronomy (miscellaneous),General Mathematics,Chemistry (miscellaneous),Computer Science (miscellaneous)

Link

https://www.mdpi.com/2073-8994/12/11/1822/pdf

Reference88 articles.

1. Double mode surveillance system based on remote audio/video signals acquisition

2. Using One-Class SVMs and Wavelets for Audio Surveillance

3. Quantifying human exposure to air pollution—Moving from static monitoring to spatio-temporally resolved personal exposure assessment

Cited by 33 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Deep transfer learning-based bird species classification using mel spectrogram images;PLOS ONE;2024-08-12

2. An intelligent epistemological tool for audiovisual analysis and mediation of video art archive;Journal of Cultural Heritage;2024-07

3. A novel approach to build a low complexity smart sound recognition system for domestic environment;Applied Acoustics;2024-05

4. Soundscape Characterization Using Autoencoders and Unsupervised Learning;Sensors;2024-04-18

5. Artificial intelligence framework for heart disease classification from audio signals;Scientific Reports;2024-02-07