Hindi and Punjabi Continuous Speech Recognition Using CNSVM-Reference-Cited by-同舟云学术

Hindi and Punjabi Continuous Speech Recognition Using CNSVM

Published:2019-10 Issue:4 Volume:11 Page:1-15
ISSN:1937-965X
Container-title:International Journal of Advanced Pervasive and Ubiquitous Computing
language:en
Short-container-title:

Author:

Passricha Vishal¹,Singhal Shubhanshi²

Affiliation:

1. National Institute of Technology, Kurukshetra, India

2. Technical Education and Research Integrated Institute, Kurukshetra, India

Abstract

CNNs are playing a vital role in the field of automatic speech recognition. Most CNNs employ a softmax activation layer to minimize cross-entropy loss. This layer generates the posterior probability in object classification tasks. SVMs are also offering promising results in the field of ASR. In this article, two different approaches: CNNs and SVMs, are combined together to propose a new hybrid architecture. This model replaces the softmax layer, i.e. the last layer of CNN by SVMs to effectively deal with high dimensional features. This model should be interpreted as a special form of structured SVM and named the Convolutional Neural SVM. (CNSVM). CNSVMs incorporate the characteristics of both models which CNNs learn features from the speech signal and SVMs classify these features into corresponding text. The parameters of CNNs and SVMs are trained jointly using a sequence level max-margin and sMBR criterion. The performance achieved by CNSVM on Hindi and Punjabi speech corpus for word error rate is 13.43% and 15.86%, respectively, which is a significant improvement on CNNs.

Publisher

IGI Global

Subject

General Medicine

Reference53 articles.

1. Abdel-Hamid, O., Deng, L., & Yu, D. (2013). Exploring convolutional neural network structures and optimization techniques for speech recognition. Paper presented at the Interspeech. Academic Press.

2. Convolutional Neural Networks for Speech Recognition

3. Applying Convolutional Neural Networks concepts to hybrid NN-HMM model for speech recognition

4. Gammatone wavelet Cepstral Coefficients for robust speech recognition

5. Agarap, A. F. (2017). An Architecture Combining Convolutional Neural Network (CNN) and Support Vector Machine (SVM) for Image Classification.