A Hybrid Time-Distributed Deep Neural Architecture for Speech Emotion Recognition-Reference-Cited by-同舟云学术

A Hybrid Time-Distributed Deep Neural Architecture for Speech Emotion Recognition

Published:2022-05-12 Issue:06 Volume:32 Page:
ISSN:0129-0657
Container-title:International Journal of Neural Systems
language:en
Short-container-title:Int. J. Neur. Syst.

Author:

De Lope Javier¹,Graña Manuel²

Affiliation:

1. Department of Artificial Intelligence, Universidad Politécnica de Madrid (UPM), Madrid, Spain

2. Computational Intelligence Group, University of the Basque Country (UPV), San Sebastian, Spain

Abstract

In recent years, speech emotion recognition (SER) has emerged as one of the most active human–machine interaction research areas. Innovative electronic devices, services and applications are increasingly aiming to check the user emotional state either to issue alerts under some predefined conditions or to adapt the system responses to the user emotions. Voice expression is a very rich and noninvasive source of information for emotion assessment. This paper presents a novel SER approach based on that is a hybrid of a time-distributed convolutional neural network (TD-CNN) and a long short-term memory (LSTM) network. Mel-frequency log-power spectrograms (MFLPSs) extracted from audio recordings are parsed by a sliding window that selects the input for the TD-CNN. The TD-CNN transforms the input image data into a sequence of high-level features that are feed to the LSTM, which carries out the overall signal interpretation. In order to reduce overfitting, the MFLPS representation allows innovative image data augmentation techniques that have no immediate equivalent on the original audio signal. Validation of the proposed hybrid architecture achieves an average recognition accuracy of 73.98% on the most widely and hardest publicly distributed database for SER benchmarking. A permutation test confirms that this result is significantly different from random classification ([Formula: see text]). The proposed architecture outperforms state-of-the-art deep learning models as well as conventional machine learning techniques evaluated on the same database trying to identify the same number of emotions.

Funder

FEDER

European Union's Horizon 2020

Publisher

World Scientific Pub Co Pte Ltd

Subject

Computer Networks and Communications,General Medicine

Link

https://www.worldscientific.com/doi/pdf/10.1142/S0129065722500241

Reference98 articles.

1. The Nature of Emotions

2. The representation and plasticity of body emotion expression

3. LieToMe: An Ensemble Approach for Deception Detection from Facial Cues

Cited by 12 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Systematic Review of Emotion Detection with Computer Vision and Deep Learning;Sensors;2024-05-28

2. Spatio-Temporal Image-Based Encoded Atlases for EEG Emotion Recognition;International Journal of Neural Systems;2024-03-27

3. Optimal Electrodermal Activity Segment for Enhanced Emotion Recognition Using Spectrogram-Based Feature Extraction and Machine Learning;International Journal of Neural Systems;2024-03-21

4. Multiple Classification of Brain MRI Autism Spectrum Disorder by Age and Gender Using Deep Learning;Journal of Medical Systems;2024-01-22

5. Cultural Differences in the Assessment of Synthetic Voices;International Journal of Neural Systems;2024-01-20