Multi-Modal Fusion Emotion Recognition Method of Speech Expression Based on Deep Learning-Reference-Cited by-同舟云学术

Multi-Modal Fusion Emotion Recognition Method of Speech Expression Based on Deep Learning

Published:2021-07-09 Issue: Volume:15 Page:
ISSN:1662-5218
Container-title:Frontiers in Neurorobotics
language:
Short-container-title:Front. Neurorobot.

Author:

Liu Dong,Wang Zhiyong,Wang Lifeng,Chen Longxi

Abstract

The redundant information, noise data generated in the process of single-modal feature extraction, and traditional learning algorithms are difficult to obtain ideal recognition performance. A multi-modal fusion emotion recognition method for speech expressions based on deep learning is proposed. Firstly, the corresponding feature extraction methods are set up for different single modalities. Among them, the voice uses the convolutional neural network-long and short term memory (CNN-LSTM) network, and the facial expression in the video uses the Inception-Res Net-v2 network to extract the feature data. Then, long and short term memory (LSTM) is used to capture the correlation between different modalities and within the modalities. After the feature selection process of the chi-square test, the single modalities are spliced to obtain a unified fusion feature. Finally, the fusion data features output by LSTM are used as the input of the classifier LIBSVM to realize the final emotion recognition. The experimental results show that the recognition accuracy of the proposed method on the MOSI and MELD datasets are 87.56 and 90.06%, respectively, which are better than other comparison methods. It has laid a certain theoretical foundation for the application of multimodal fusion in emotion recognition.

Publisher

Frontiers Media SA

Subject

Artificial Intelligence,Biomedical Engineering

Reference35 articles.

1. An appraisal on speech and emotion recognition technologies based on machine learning;Andy;Int. J. Automot. Technol.,2020

2. Facial expression synthesis using vowel recognition for synthesized speech;Asada;Artif. Life Robot.,2020

3. Human emotional state assessment based on a video portrayal;Barabanschikov;Exp. Psychol.,2020

4. Multimodal biometric recognition: fusion of modified adaptive bilinear interpolation data samples of face and signature using local binary pattern features;Bc;Int. J. Eng. Adv. Technol.,2020

5. Modeling human age-associated increase in Gadd45γ expression leads to spatial recognition memory impairments in young adult mice;Brito;Neurobiol. Aging,2020

Cited by 31 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Attention-based acoustic feature fusion network for depression detection;Neurocomputing;2024-10

2. A Two-Stage Multi-Modal Multi-Label Emotion Recognition Decision System Based on GCN;International Journal of Decision Support System Technology;2024-08-16

3. Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments;Frontiers in Plant Science;2024-07-25

4. A survey of multimodal hybrid deep learning for computer vision: Architectures, applications, trends, and challenges;Information Fusion;2024-05

5. Multimodal Emotion Recognition with Deep Learning: Advancements, challenges, and future directions;Information Fusion;2024-05