Lightweight deep convolutional neural network for background sound classification in speech signals-Reference-Cited by-同舟云学术

Lightweight deep convolutional neural network for background sound classification in speech signals

Published:2022-04 Issue:4 Volume:151 Page:2773-2786
ISSN:0001-4966
Container-title:The Journal of the Acoustical Society of America
language:en
Short-container-title:The Journal of the Acoustical Society of America

Author:

Dayal Aveen¹,Yeduri Sreenivasa Reddy¹^ORCID,Koduru Balu Harshavardan¹,Jaiswal Rahul Kumar¹,Soumya J.²,Srinivas M. B.²,Pandey Om Jee³,Cenkeramaddi Linga Reddy¹^ORCID

Affiliation:

1. Department of ICT, University of Agder, Grimstad 4879, Norway

2. Birla Institute of Technology and Science-Pilani, Hyderabad, India

3. Department of Electronics Engineering, IIT BHU Varanasi, Varanasi 221005, India

Abstract

Recognizing background information in human speech signals is a task that is extremely useful in a wide range of practical applications, and many articles on background sound classification have been published. It has not, however, been addressed with background embedded in real-world human speech signals. Thus, this work proposes a lightweight deep convolutional neural network (CNN) in conjunction with spectrograms for an efficient background sound classification with practical human speech signals. The proposed model classifies 11 different background sounds such as airplane, airport, babble, car, drone, exhibition, helicopter, restaurant, station, street, and train sounds embedded in human speech signals. The proposed deep CNN model consists of four convolution layers, four max-pooling layers, and one fully connected layer. The model is tested on human speech signals with varying signal-to-noise ratios (SNRs). Based on the results, the proposed deep CNN model utilizing spectrograms achieves an overall background sound classification accuracy of 95.2% using the human speech signals with a wide range of SNRs. It is also observed that the proposed model outperforms the benchmark models in terms of both accuracy and inference time when evaluated on edge computing devices.

Funder

Research Council of Norway

Publisher

Acoustical Society of America (ASA)

Subject

Acoustics and Ultrasonics,Arts and Humanities (miscellaneous)

Link

https://asa.scitation.org/doi/pdf/10.1121/10.0010257

Reference55 articles.

1. End-to-end environmental sound classification using a 1D convolutional neural network

2. Environmental sound classification using optimum allocation sampling based empirical mode decomposition

3. Acoustic Scene Classification: Classifying environments from the sounds they produce

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Survey on Deep Learning Based Forest Environment Sound Classification at the Edge;ACM Computing Surveys;2023-10-05

2. An acoustic tracking model based on deep learning using two hydrophones and its reverberation transfer hypothesis, applied to whale tracking;Frontiers in Marine Science;2023-08-01

3. Estimation of number of unmanned aerial vehicles in a scene utilizing acoustic signatures and machine learning;The Journal of the Acoustical Society of America;2023-07-01

4. Convolutional neural network reveals frequency content of medio-lateral COM body sway to be highly predictive of Parkinson’s disease;2023-05-30

5. Environment Knowledge-Driven Generic Models to Detect Coughs From Audio Recordings;IEEE Open Journal of Engineering in Medicine and Biology;2023