Speech Models Training Technologies Comparison Using Word Error Rate-Reference-Cited by-同舟云学术

Speech Models Training Technologies Comparison Using Word Error Rate

Published:2023-05-10 Issue:1 Volume:8 Page:74-80
ISSN:2524-0382
Container-title:Advances in Cyber-Physical Systems
language:
Short-container-title:ACPS

Author:

Yakubovskyi Roman, ,Morozov Yuriy

Abstract

The main purpose of this work is to analyze and compare several technologies used for training speech models, including traditional approaches as Hidden Markov Models (HMMs) and more recent methods as Deep Neural Networks (DNNs). The technologies have been explained and compared using word error rate metric based on the input of 1000 words by a user with 15 decibel background noise. Word error rate metric has been ex- plained and calculated. Potential replacements for com- pared technologies have been provided, including: Atten- tion-based, Generative, Sparse and Quantum-inspired models. Pros and cons of those techniques as a potential replacement have been analyzed and listed. Data analyzing tools and methods have been explained and most common datasets used for HMM and DNN technologies have been described. Real life usage examples of both methods have been provided and systems based on them have been ana- lyzed.

Publisher

Lviv Polytechnic National University

Subject

General Medicine

Reference10 articles.

1. Borovets D., Pavych T., Paramud Y., (2021). Computer System for Converting Gestures to Text and Audio Mes- sages. Advances in Cyber-Physical Systems. vol. 6, num. 2. Pp. 90-97. DOI: https://doi.org/10.23939/acps2021.02.090

2. Emiru E. D., Li Y., Xiong S., Fesseha A., (2019). Speech recognition system based on deep neural network acous- tic modeling for low resourced language-Amharic. ICTCE '19: Proceedings of the 3rd International Confer- ence on Telecommunications and Communication Engi- neering. [Online]. Pp. 141-145. DOI: https://dl.acm.org/doi/10.1145/3369555.3369564#sec- terms

3. Tanaka T., Masumura R., Moriya T., Oba T., Aono Y., (2019). A Joint End-to-End and DNN-HMM Hybrid Automatic Speech Recognition System with Transferring Sharable Knowledge. NTT Media Intelligence Laborato- ries, NTT Corporation. [Online]. Pp. 2210-2214. DOI: http://dx.doi.org/10.21437/Interspeech.2019-226

4. Shanin I., (2019). Emotion Recognition based on Third- Order Circular Suprasegmental Hidden Markov Model. 2019 IEEE Jordan International Joint Conference on Electrical Engineering and Information Technology (JEEIT). [Online]. Pp. 800-805. DOI: https://doi.org/10.1109/ICASSP.2019.8683172

5. Dutta A., Ashishkumar G., Rama Rao C. V., (2021). Performance analysis of ASR system in hybrid DNN- HMM framework using a PWL euclidean activation function. Frontiers of Computer Science. [Online]. Pp. 2095-2236. DOI: https://doi.org/10.1007/s11704-020-9419-z

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Arabic Automatic Speech Recognition: Challenges and Progress;Speech Communication;2024-09