End-to-End Modeling and Transfer Learning for Audiovisual Emotion Recognition in-the-Wild-Reference-Cited by-同舟云学术

End-to-End Modeling and Transfer Learning for Audiovisual Emotion Recognition in-the-Wild

Published:2022-01-27 Issue:2 Volume:6 Page:11
ISSN:2414-4088
Container-title:Multimodal Technologies and Interaction
language:en
Short-container-title:MTI

Author:

Dresvyanskiy Denis,Ryumina Elena^ORCID,Kaya Heysem^ORCID,Markitantov Maxim,Karpov Alexey^ORCID,Minker Wolfgang

Abstract

As emotions play a central role in human communication, automatic emotion recognition has attracted increasing attention in the last two decades. While multimodal systems enjoy high performances on lab-controlled data, they are still far from providing ecological validity on non-lab-controlled, namely “in-the-wild” data. This work investigates audiovisual deep learning approaches to emotion recognition in in-the-wild problem. Inspired by the outstanding performance of end-to-end and transfer learning techniques, we explored the effectiveness of architectures in which a modality-specific Convolutional Neural Network (CNN) is followed by a Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) using the AffWild2 dataset under the Affective Behavior Analysis in-the-Wild (ABAW) challenge protocol. We deployed unimodal end-to-end and transfer learning approaches within a multimodal fusion system, which generated final predictions using a weighted score fusion scheme. Exploiting the proposed deep-learning-based multimodal system, we reached a test set challenge performance measure of 48.1% on the ABAW 2020 Facial Expressions challenge, which advances the first-runner-up performance.

Funder

Russian Foundation for Basic Research

Russian state research

Publisher

MDPI AG

Subject

Computer Networks and Communications,Computer Science Applications,Human-Computer Interaction,Neuroscience (miscellaneous)

Link

https://www.mdpi.com/2414-4088/6/2/11/pdf

Reference95 articles.

1. Affective Computing;Picard,2000

2. Call Redistribution for a Call Center Based on Speech Emotion Recognition

3. An Emotion Recognition Model Based on Facial Recognition in Virtual Learning Environment

Cited by 17 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Exploring contactless techniques in multimodal emotion recognition: insights into diverse applications, challenges, solutions, and prospects;Multimedia Systems;2024-04-06

2. Advances in Facial Expression Recognition: A Survey of Methods, Benchmarks, Models, and Datasets;Information;2024-02-28

3. Deep Learning-Based Automatic Speech and Emotion Recognition for Students with Disabilities: A Review;Applied Intelligence and Informatics;2024

4. Audio-Visual Self-Supervised Representation Learning: A Survey;2024

5. Facial Emotion Recognition in-the-Wild Using Deep Neural Networks: A Comprehensive Review;SN Computer Science;2023-12-13