Multi-rate modulation encoding via unsupervised learning for audio event detection-Reference-Cited by-同舟云学术

Multi-rate modulation encoding via unsupervised learning for audio event detection

Published:2024-04-01 Issue:1 Volume:2024 Page:
ISSN:1687-4722
Container-title:EURASIP Journal on Audio, Speech, and Music Processing
language:en
Short-container-title:J AUDIO SPEECH MUSIC PROC.

Author:

Kothinti Sandeep Reddy,Elhilali Mounya^ORCID

Abstract

AbstractTechnologies in healthcare, smart homes, security, ecology, and entertainment all deploy audio event detection (AED) in order to detect sound events in an audio recording. Effective AED techniques rely heavily on supervised or semi-supervised models to capture the wide range of dynamics spanned by sound events in order to achieve temporally precise boundaries and accurate event classification. These methods require extensive collections of labeled or weakly labeled in-domain data, which is costly and labor-intensive. Importantly, these approaches do not fully leverage the inherent variability and range of dynamics across sound events, aspects that can be effectively identified through unsupervised methods. The present work proposes an approach based on multi-rate autoencoders that are pretrained in an unsupervised way to leverage unlabeled audio data and ultimately learn the rich temporal dynamics inherent in natural sound events. This approach utilizes parallel autoencoders that achieve decompositions of the modulation spectrum along different bands. In addition, we introduce a rate-selective temporal contrastive loss to align the training objective with event detection metrics. Optimizing the configuration of multi-rate encoders and the temporal contrastive loss leads to notable improvements in domestic sound event detection in the context of the DCASE challenge.

Funder

Office of Naval Research Global

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1186/s13636-024-00339-5.pdf

Reference43 articles.

1. Y. Zigel, D. Litvak, I. Gannot, A method for automatic fall detection of elderly people using floor vibrations and sound - Proof of concept on human mimicking doll falls. IEEE Trans. Biomed. Eng. 56(12), 2858–2867 (2009). https://doi.org/10.1109/TBME.2009.2030171

2. Q. Jin, P.F. Schulam, S. Rawat, S. Burger, D. Ding, F. Metze, Event-based video retrieval using audio, in 13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012. (International Speech Communication Association, Lyon, 2012). https://doi.org/10.21437/Interspeech.2012-556

3. A.O. Eren, M. Sert, in 2020 IEEE International Symposium on Multimedia (ISM). Audio captioning based on combined audio and semantic embeddings. (2020), pp. 41–48. https://doi.org/10.1109/ISM.2020.00014

4. H. Phan, P. Koch, F. Katzberg, M. Maass, R. Mazur, I. McLoughlin, A. Mertins, 25th European Signal Processing Conference, EUSIPCO 2017 2017-January. What makes audio event detection harder than classification? (2017), pp. 2739–2743. https://doi.org/10.23919/EUSIPCO.2017.8081709

5. H. Phan, T.N.T. Nguyen, P. Koch, A. Mertins, in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Polyphonic audio event detection: Multi-label or multi-class multi-task classification problem? (IEEE, 2022), pp. 8877–8881.https://doi.org/10.1109/ICASSP43922.2022.9746402

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Neural Cough Counter: A Novel Deep Learning Approach for Cough Detection and Monitoring;IEEE Access;2024