Gender and Age Estimation Methods Based on Speech Using Deep Neural Networks-Reference-Cited by-同舟云学术

Gender and Age Estimation Methods Based on Speech Using Deep Neural Networks

Published:2021-07-13 Issue:14 Volume:21 Page:4785
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Kwasny Damian^ORCID,Hemmerling Daria^ORCID

Abstract

The speech signal contains a vast spectrum of information about the speaker such as speakers’ gender, age, accent, or health state. In this paper, we explored different approaches to automatic speaker’s gender classification and age estimation system using speech signals. We applied various Deep Neural Network-based embedder architectures such as x-vector and d-vector to age estimation and gender classification tasks. Furthermore, we have applied a transfer learning-based training scheme with pre-training the embedder network for a speaker recognition task using the Vox-Celeb1 dataset and then fine-tuning it for the joint age estimation and gender classification task. The best performing system achieves new state-of-the-art results on the age estimation task using popular TIMIT dataset with a mean absolute error (MAE) of 5.12 years for male and 5.29 years for female speakers and a root-mean square error (RMSE) of 7.24 and 8.12 years for male and female speakers, respectively, and an overall gender recognition accuracy of 99.60%.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/21/14/4785/pdf

Reference37 articles.

1. Paralinguistics in speech and language—State-of-the-art and the challenge

2. Acoustic analysis assessment in speech pathology detection

3. Techmohttps://www.techmo.pl

4. Age Estimation in Short Speech Utterances Based on LSTM Recurrent Neural Networks

Cited by 29 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Voice Privacy Using Time-Scale and Pitch Modification;SN Computer Science;2024-01-27

2. Speaker age and gender recognition using 1D and 2D convolutional neural networks;Neural Computing and Applications;2023-11-28

3. Evaluating the Performance of wav2vec Embedding for Parkinson's Disease Detection;Measurement Science Review;2023-11-17

4. AgeNet-AT: An End-to-End Model for Robust Joint Speaker Age Estimation and Gender Recognition Based on Attention Mechanism and Titanet;2023 13th International Conference on Computer and Knowledge Engineering (ICCKE);2023-11-01

5. Coarse-Age Loss: A New Training Method Using Coarse-Age Labeled Data for Speaker Age Estimation;2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC);2023-10-31