Swahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition-Reference-Cited by-同舟云学术

Swahili Speech Dataset Development and Improved Pre-training Method for Spoken Digit Recognition

Published:2023-07-20 Issue:7 Volume:22 Page:1-24
ISSN:2375-4699
Container-title:ACM Transactions on Asian and Low-Resource Language Information Processing
language:en
Short-container-title:ACM Trans. Asian Low-Resour. Lang. Inf. Process.

Author:

Kivaisi Alexander R.¹^ORCID,Zhao Qingjie¹^ORCID,Mbelwa Jimmy T.²^ORCID

Affiliation:

1. Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science and Technology, Beijing Institute of Technology

2. Department of Computer Science and Engineering, University of Dar es salaam

Abstract

Speech dataset is an essential component in building commercial speech applications. However, low-resource languages such as Swahili lack such a resource that is vital for spoken digit recognition. For languages where such resources exist, they are usually insufficient. Thus, pre-training methods have been used with external resources to improve continuous speech recognition. However, to the best of our knowledge, no study has investigated the effect of pre-training methods specifically for spoken digit recognition. This study aimed at addressing these problems. First, we developed a Swahili spoken digit dataset for Swahili spoken digit recognition. Then, we investigated the effect of cross-lingual and multi-lingual pre-training methods on spoken digit recognition. Finally, we proposed an effective language-independent pre-training method for spoken digit recognition. The proposed method has the advantage of incorporating target language data during the pre-training stage that leads to an optimal solution when using less training data. Experiments on Swahili (being developed), English, and Gujarati datasets show that our method achieves better performance compared with all the baselines listed in this study.

Funder

China Scholarship Council

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3597494

Reference41 articles.

1. Martín Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Mane Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Viegas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2016. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv preprint arXiv:1603.04467.

2. Massively multilingual adversarial speech recognition;Adams Oliver;NAACL HLT 2019 - 2019 Conf. North Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. - Proc. Conf,2019

3. Database development and automatic speech recognition of isolated Pashto spoken digits using MFCC and K-NN

4. A. de Andrade Bresolin, A. D. D. Neto, and P. J. Alsina. 2008. Digit recognition using wavelet and SVM in Brazilian Portuguese. In Proceedings of the 2008 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1545–1548.

5. A tutorial on onset detection in music signals

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Improved mini-batch multiple augmentation for low-resource spoken word recognition;Expert Systems with Applications;2024-10

2. A Multifaceted Feature Extraction Approach for Noise-Robust Punjabi Spoken Digit Recognition System Under Low-Resource Conditions;2024 11th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO);2024-03-14

3. The First Swahili Language Scene Text Detection and Recognition Dataset;Lecture Notes in Computer Science;2024