Dynamic out-of-vocabulary word registration to language model for speech recognition-Reference-Cited by-同舟云学术

Dynamic out-of-vocabulary word registration to language model for speech recognition

Published:2021-01-25 Issue:1 Volume:2021 Page:
ISSN:1687-4722
Container-title:EURASIP Journal on Audio, Speech, and Music Processing
language:en
Short-container-title:J AUDIO SPEECH MUSIC PROC.

Author:

Kitaoka Norihide^ORCID,Chen Bohan,Obashi Yuya

Abstract

AbstractWe propose a method of dynamically registering out-of-vocabulary (OOV) words by assigning the pronunciations of these words to pre-inserted OOV tokens, editing the pronunciations of the tokens. To do this, we add OOV tokens to an additional, partial copy of our corpus, either randomly or to part-of-speech (POS) tags in the selected utterances, when training the language model (LM) for speech recognition. This results in an LM containing OOV tokens, to which we can assign pronunciations. We also investigate the impact of acoustic complexity and the “natural” occurrence frequency of OOV words on the recognition of registered OOV words. The proposed OOV word registration method is evaluated using two modern automatic speech recognition (ASR) systems, Julius and Kaldi, using DNN-HMM acoustic models and N-gram language models (plus an additional evaluation using RNN re-scoring with Kaldi). Our experimental results show that when using the proposed OOV registration method, modern ASR systems can recognize OOV words without re-training the language model, that the acoustic complexity of OOV words affects OOV recognition, and that differences between the “natural” and the assigned occurrence frequencies of OOV words have little impact on the final recognition results.

Funder

Japan Society for the Promotion of Science

Publisher

Springer Science and Business Media LLC

Subject

Electrical and Electronic Engineering,Acoustics and Ultrasonics

Link

http://link.springer.com/content/pdf/10.1186/s13636-020-00193-1.pdf

Reference21 articles.

1. I. Bazzi, J. R. Glass, in ICSLP-2000. Modeling out-of-vocabulary words for robust speech recognition (ISCA, 2000), pp. 401–404.

2. I. Bazzi, J. R. Glass, in ICSLP-2002. A multi-class approach for modelling out-of-vocabulary words (ISCA, 2002), pp. 1613–1616.

3. M. Creutz, T. Hirsimaki, M. Kurimo, A. Puurula, J. Pylkkonen, V. Siivola, M. Varjokallio, E. Arisoy, M. Saraclar, A. Stolcke, in NAACL-HLT 2007. Analysis of morph-based speech recognition and the modeling of out-of-vocabulary words across languages (ACL, 2007), pp. 380–397.

4. H. Sun, G. Zhang, M. Xu, in EUROSPEECH2003. Using word confidence measure for OOV words detection in a spontaneous spoken dialog system (ISCA, 2003), pp. 2713–2716.

5. B. Lecouteux, G. Linares, B. Favre, in EUROSPEECH2003. Using word confidence measure for OOV words detection in a spontaneous spoken dialog system (ISCA, 2003), pp. 2713–2716.

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Initial decoding with minimally augmented language model for improved lattice rescoring in low resource ASR;Sādhanā;2024-05-21

2. A comprehensive survey on automatic speech recognition using neural networks;Multimedia Tools and Applications;2023-08-15

3. Do smart speaker skills support diverse audiences?;Pervasive and Mobile Computing;2022-12

4. Adaptive Discounting of Implicit Language Models in RNN-Transducers;ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP);2022-05-23