Data Augmentation Using Spectral Warping for Low Resource Children ASR-Reference-Cited by-同舟云学术

Data Augmentation Using Spectral Warping for Low Resource Children ASR

Published:2022-11-08 Issue:12 Volume:94 Page:1507-1513
ISSN:1939-8018
Container-title:Journal of Signal Processing Systems
language:en
Short-container-title:J Sign Process Syst

Author:

Kathania Hemant Kumar^ORCID,Kadyan Viredner,Kadiri Sudarsana Reddy^ORCID,Kurimo Mikko

Abstract

AbstractIn low resource children automatic speech recognition (ASR) the performance is degraded due to limited acoustic and speaker variability available in small datasets. In this paper, we propose a spectral warping based data augmentation method to capture more acoustic and speaker variability. This is carried out by warping the linear prediction (LP) spectra computed from speech data. The warped LP spectra computed in a frame-based manner are used with the corresponding LP residuals to synthesize speech to capture more variability. The proposed augmentation method is shown to improve the ASR system performance over the baseline system. We have compared the proposed method with four well-known data augmentation methods: pitch scaling, speaking rate, SpecAug and vocal tract length perturbation (VTLP), and found that the proposed method performs the best. Further, we have combined the proposed method with these existing data augmentation methods to improve the ASR system performance even more. The combined system consisting of the original data, VTLP, SpecAug and the proposed spectral warping method gave the best performance by a relative word error rate reduction of 32.13% and 10.51% over the baseline system for Punjabi children and TLT-school corpus, respectively. The proposed spectral warping method is publicly available at https://github.com/kathania/Spectral-Warping.

Funder

Aalto University

Publisher

Springer Science and Business Media LLC

Subject

Hardware and Architecture,Modeling and Simulation,Information Systems,Signal Processing,Theoretical Computer Science,Control and Systems Engineering

Link

https://link.springer.com/content/pdf/10.1007/s11265-022-01820-0.pdf

Reference31 articles.

1. Evermann, G., Chan, H. Y., Gales, M. J., Hain, T., Liu, X., Mrva, D., Wang, L., & Woodland, P. C. (2004). Development of the 2003 CU-HTK conversational telephone speech transcription system. In 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing (vol. 1, p. 249). IEEE.

2. Kanda, N., Takeda, R., & Obuchi, Y. (2013). Elastic spectral distortion for low resource speech recognition with deep neural networks. In 2013 IEEE Workshop on Automatic Speech Recognition and Understanding (pp. 309–314). IEEE.

3. Hu, H., Tan, T., & Qian, Y. (2018). Generative adversarial networks based data augmentation for noise robust speech recognition. In Proceedings - ICASSP (pp. 5044–5048).

4. Qian, Y., Hu, H., & Tan, T. (2019). Data augmentation using generative adversarial networks for robust speech recognition. Speech Communication, 114, 1–9.

5. Gales, M. J., Ragni, A., AlDamarki, H., & Gautier, C. (2009). Support vector machines for noise robust ASR. In 2009 IEEE Workshop on Automatic Speech Recognition & Understanding (pp. 205–210). IEEE.

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Effect of Speech Modification on Wav2Vec2 Models for Children Speech Recognition;2024 International Conference on Signal Processing and Communications (SPCOM);2024-07-01

2. In-Domain Data Augmentation to Enhance Severity Level Classification of Dysarthria from Speech;2024 International Conference on Signal Processing and Communications (SPCOM);2024-07-01

3. ChildAugment: Data augmentation methods for zero-resource children's speaker verification;The Journal of the Acoustical Society of America;2024-03-01

4. Improved Vocal Tract Length Perturbation for Improving Child Speech Emotion Recognition;2023 International Conference on Sensing, Measurement & Data Analytics in the era of Artificial Intelligence (ICSMD);2023-11-02

5. Comparison of Data Augmentation Techniques on Filipino ASR for Children’s Speech;2023 International Conference on Speech Technology and Human-Computer Dialogue (SpeD);2023-10-25