The Power-Normalized Cepstral Coefficient (PNCC) for convolutional neural networks-based robust speech command recognition
-
Published:2023-09-01
Issue:1
Volume:2596
Page:012021
-
ISSN:1742-6588
-
Container-title:Journal of Physics: Conference Series
-
language:
-
Short-container-title:J. Phys.: Conf. Ser.
Author:
Iswanto B H,Hafizhahullah H,Pardede H F,Zahra A
Abstract
Abstract
While implementations of speech recognition grow rapidly in recent years and are slowly being integrated into our daily devices, the problem of noise robustness is still a challenging task, even with the recent advancement of deep learning technologies for speech recognition. The presence of noise may cause a mismatch between training, which is performed in clean conditions, and noisy testing conditions. This paper proposes a method to extract features for speech recognition by employing features derived under the power law scale, i.e., the Power-Normalized Cepstral Coefficient (PNCC). The power-law can provide better compression in low-energy regions so that it is not sensitive when the speech signal is distorted by noise. The features are implemented on speech recognition based on Convolutional Neural Networks (CNNs). The experiments were carried out by TensorFlow’s Speech Command Dataset mixed with various signal-to-noise ratio to evaluate the method. The experimental findings indicate that the accuracy ranges from 81% to 86%.
Subject
Computer Science Applications,History,Education
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献