An Effective Speaker Clustering Method using UBM and Ultra-Short Training Utterances-Reference-Cited by-同舟云学术

An Effective Speaker Clustering Method using UBM and Ultra-Short Training Utterances

Published:2016-03-01 Issue:1 Volume:41 Page:107-118
ISSN:2300-262X
Container-title:Archives of Acoustics
language:
Short-container-title:

Author:

Hossa Robert,Makowski Ryszard

Abstract

Abstract The same speech sounds (phones) produced by different speakers can sometimes exhibit significant differences. Therefore, it is essential to use algorithms compensating these differences in ASR systems. Speaker clustering is an attractive solution to the compensation problem, as it does not require long utterances or high computational effort at the recognition stage. The report proposes a clustering method based solely on adaptation of UBM model weights. This solution has turned out to be effective even when using a very short utterance. The obtained improvement of frame recognition quality measured by means of frame error rate is over 5%. It is noteworthy that this improvement concerns all vowels, even though the clustering discussed in this report was based only on the phoneme a. This indicates a strong correlation between the articulation of different vowels, which is probably related to the size of the vocal tract.

Publisher

Walter de Gruyter GmbH

Subject

Acoustics and Ultrasonics

Link

https://www.degruyter.com/view/j/aoa.2016.41.issue-1/aoa-2016-0011/aoa-2016-0011.pdf

Reference8 articles.

1. Partially Supervised Speaker Clustering IEEE Transaction on Pattern Analysis and Machine;TangH;Intelligence,2012

2. Speaker clustering for speech recognition using vocal track parameters;NaitoM;Speech Communication,2002

3. Distance measures for signal processing and pattern recognition Processing;BassevilleM;Signal,1989

4. Singing speaker clustering based on subspace learning in the GMM mean supervector space;MehrabaniM;Speech Communication,2013

5. Minimum Hellinger distance estimation for finite Poisson regression models and its applications;LuZ;Biometrics,2003

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Amplitude spectrum correction to improve speech signal classification quality;International Journal of Electronics and Telecommunications;2024-07-25

2. Speaker Model Clustering to Construct Background Models for Speaker Verification;Archives of Acoustics;2017-03-01