Multimodal Emotion Detection via Attention-Based Fusion of Extracted Facial and Speech Features-Reference-Cited by-同舟云学术

Multimodal Emotion Detection via Attention-Based Fusion of Extracted Facial and Speech Features

Published:2023-06-09 Issue:12 Volume:23 Page:5475
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Mamieva Dilnoza¹,Abdusalomov Akmalbek Bobomirzaevich¹^ORCID,Kutlimuratov Alpamis²,Muminov Bahodir³,Whangbo Taeg Keun¹

Affiliation:

1. Department of Computer Engineering, Gachon University, Seongnam-si 13120, Republic of Korea

2. Department of AI. Software, Gachon University, Seongnam-si 13120, Republic of Korea

3. Department of Artificial Intelligence, Tashkent State University of Economics, Tashkent 100066, Uzbekistan

Abstract

Methods for detecting emotions that employ many modalities at the same time have been found to be more accurate and resilient than those that rely on a single sense. This is due to the fact that sentiments may be conveyed in a wide range of modalities, each of which offers a different and complementary window into the thoughts and emotions of the speaker. In this way, a more complete picture of a person’s emotional state may emerge through the fusion and analysis of data from several modalities. The research suggests a new attention-based approach to multimodal emotion recognition. This technique integrates facial and speech features that have been extracted by independent encoders in order to pick the aspects that are the most informative. It increases the system’s accuracy by processing speech and facial features of various sizes and focuses on the most useful bits of input. A more comprehensive representation of facial expressions is extracted by the use of both low- and high-level facial features. These modalities are combined using a fusion network to create a multimodal feature vector which is then fed to a classification layer for emotion recognition. The developed system is evaluated on two datasets, IEMOCAP and CMU-MOSEI, and shows superior performance compared to existing models, achieving a weighted accuracy WA of 74.6% and an F1 score of 66.1% on the IEMOCAP dataset and a WA of 80.7% and F1 score of 73.7% on the CMU-MOSEI dataset.

Funder

GRRC program of Gyeonggi province

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/23/12/5475/pdf

Reference53 articles.

1. Biele, C., Kacprzyk, J., Kopeć, W., Owsiński, J.W., Romanowski, A., and Sikorski, M. (2022). Digital Interaction and Machine Intelligence, 9th Machine Intelligence and Digital Interaction Conference, Warsaw, Poland, 9–10 December 2021, Springer. Lecture Notes in Networks and Systems.

2. A systematic survey on multimodal emotion recognition using learning algorithms;Ahmed;Intell. Syst. Appl.,2023

3. Gu, X., Shen, Y., and Xu, J. (2021, January 18–21). Multimodal Emotion Recognition in Deep Learning:a Survey. Proceedings of the 2021 International Conference on Culture-Oriented Science & Technology (ICCST), Beijing, China.