Speaker Identification Using Empirical Mode Decomposition-Based Voice Activity Detection Algorithm under Realistic Conditions-Reference-Cited by-同舟云学术

Speaker Identification Using Empirical Mode Decomposition-Based Voice Activity Detection Algorithm under Realistic Conditions

Published:2014-12-01 Issue:4 Volume:23 Page:405-421
ISSN:2191-026X
Container-title:Journal of Intelligent Systems
language:
Short-container-title:

Author:

Rudramurthy M.S.¹,Pathak Nilabh Kumar¹,Prasad V. Kamakshi²,Kumaraswamy R.³

Affiliation:

1. 1Department of IS&E, Siddaganga Institute of Technology, Tumkur 572103, Karnataka, India

2. 2School of Information Technology, JNTU, Kukatpally, Hyderabad 500085, Andhra Pradesh, India

3. 3Senior Member, IEEE, Department of EC&E, Siddaganga Institute of Technology, Tumkur 572103, Karnataka, India

Abstract

AbstractSpeaker recognition (SR) under mismatched conditions is a challenging task. Speech signal is nonlinear and nonstationary, and therefore, difficult to analyze under realistic conditions. Also, in real conditions, the nature of the noise present in speech data is not known a priori. In such cases, the performance of speaker identification (SI) or speaker verification (SV) degrades considerably under realistic conditions. Any SR system uses a voice activity detector (VAD) as the front-end subsystem of the whole system. The performance of most VADs deteriorates at the front end of the SR task or system under degraded conditions or in realistic conditions where noise plays a major role. Recently, speech data analysis and processing using Norden E. Huang’s empirical mode decomposition (EMD) combined with Hilbert transform, commonly referred to as Hilbert–Huang transform (HHT), has become an emerging trend. EMD is an a posteriori, adaptive, data analysis tool used in time domain that is widely accepted by the research community. Recently, speech data analysis and speech data processing for speech recognition and SR tasks using EMD have been increasing. EMD-based VAD has become an important adaptive subsystem of the SR system that mostly mitigates the effect of mismatch between the training and the testing phase. Recently, we have developed a VAD algorithm using a zero-frequency filter-assisted peaking resonator (ZFFPR) and EMD. In this article, the efficacy of an EMD-based VAD algorithm is studied at the front end of a text-independent language-independent SI task for the speaker’s data collected in three languages at five different places, such as home, street, laboratory, college campus, and restaurant, under realistic conditions using EDIROL-R09 HR, a 24-bit wav/MP3 recorder. The performance of this proposed SI task is compared against the traditional energy-based VAD in terms of percentage identification rate. In both cases, widely accepted Mel frequency cepstral coefficients are computed by employing frame processing (20-ms frame size and 10-ms frame shift) from the extracted voiced speech regions using the respective VAD techniques from the realistic speech utterances, and are used as a feature vector for speaker modeling using popular Gaussian mixture models. The experimental results showed that the proposed SI task with the VAD algorithm using ZFFPR and EMD at its front end performs better than the SI task with short-term energy-based VAD when used at its front end, and is somewhat encouraging.

Publisher

Walter de Gruyter GmbH

Subject

Artificial Intelligence,Information Systems,Software

Link

https://www.degruyter.com/document/doi/10.1515/jisys-2013-0089/pdf

Reference148 articles.

1. An analysis of the embedded frequency content of macroeconomic indicators and their counterparts using the Hilbert transform Bank of Finland Research Discussion;Crowley;Papers,2009

2. The quefrency alanysis of time series for echoes : cepstrum pseudo autocovariance cross - cepstrum and saphe cracking in : Proceedings of the Symposium on Time Series Analysis ed Chapter pp New York;Bogert,1963

3. Overview of speaker enhancement techniques for automatic speaker recognition in Proceedings of Fourth International Conference on Spoken Language Processing;Ortega;October,1996

4. Nonlinear evolution of water waves s view in Proceedings of the International Symposium on nd ed eds World Scientific Scotland UK;Huang;Experimental Chaos,1995

5. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences;Davis;IEEE Acoust Speech,1980

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Improved Empirical Mode Decomposition Using Optimal Recursive Averaging Noise Estimation for Speech Enhancement;Circuits, Systems, and Signal Processing;2021-06-21

2. Information-Theoretic Measures on Intrinsic Mode Function for the Individual Identification Using EEG Sensors;IEEE Sensors Journal;2015-09

3. Feature-level fusion of mental task’s brain signal for an efficient identification system;Neural Computing and Applications;2015-05-07