Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms-Reference-Cited by-同舟云学术

Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms

Published:2021-09-01 Issue:17 Volume:21 Page:5892
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Tursunov Anvarjon^ORCID,Mustaqeem ^ORCID,Choeh Joon Yeon^ORCID,Kwon Soonil^ORCID

Abstract

Speech signals are being used as a primary input source in human–computer interaction (HCI) to develop several applications, such as automatic speech recognition (ASR), speech emotion recognition (SER), gender, and age recognition. Classifying speakers according to their age and gender is a challenging task in speech processing owing to the disability of the current methods of extracting salient high-level speech features and classification models. To address these problems, we introduce a novel end-to-end age and gender recognition convolutional neural network (CNN) with a specially designed multi-attention module (MAM) from speech signals. Our proposed model uses MAM to extract spatial and temporal salient features from the input data effectively. The MAM mechanism uses a rectangular shape filter as a kernel in convolution layers and comprises two separate time and frequency attention mechanisms. The time attention branch learns to detect temporal cues, whereas the frequency attention module extracts the most relevant features to the target by focusing on the spatial frequency features. The combination of the two extracted spatial and temporal features complements one another and provide high performance in terms of age and gender classification. The proposed age and gender classification system was tested using the Common Voice and locally developed Korean speech recognition datasets. Our suggested model achieved 96%, 73%, and 76% accuracy scores for gender, age, and age-gender classification, respectively, using the Common Voice dataset. The Korean speech recognition dataset results were 97%, 97%, and 90% for gender, age, and age-gender recognition, respectively. The prediction performance of our proposed model, which was obtained in the experiments, demonstrated the superiority and robustness of the tasks regarding age, gender, and age-gender recognition from speech signals.

Funder

Ministry of Science and ICT, South Korea

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/21/17/5892/pdf

Reference52 articles.

1. Improved noisy student training for automatic speech recognition;Park;arXiv,2020

2. Deep-Net: A Lightweight CNN-Based Speech Emotion Recognition System Using Deep Frequency Features

3. Speaker age estimation using i-vectors

Cited by 49 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Konuşmacıları Kadın, Erkek ve Çocuk Olarak Sınıflandırmada Veri Artırmanın Performansa Etkisi;Iğdır Üniversitesi Fen Bilimleri Enstitüsü Dergisi;2024-09-01

2. An optimized attention based hybrid deep learning framework for automatic speaker identification from speech signals;Multimedia Tools and Applications;2024-08-23

3. Novel SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech;Circuits, Systems, and Signal Processing;2024-08-08

4. Automatic Age and Gender Recognition Using Ensemble Learning;Applied Sciences;2024-08-06

5. Gender Recognition Based on the Stacking of Different Acoustic Features;Applied Sciences;2024-07-27