Multi-Label Emotion Recognition of Korean Speech Data Using Deep Fusion Models-Reference-Cited by-同舟云学术

Multi-Label Emotion Recognition of Korean Speech Data Using Deep Fusion Models

Published:2024-08-28 Issue:17 Volume:14 Page:7604
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Park Seoin¹,Jeon Byeonghoon¹,Lee Seunghyun¹,Yoon Janghyeok¹

Affiliation:

1. Department of Industrial Engineering, Konkuk University, 120 Neungdong-ro, Gwangjin-gu, Seoul 05029, Republic of Korea

Abstract

As speech is the most natural way for humans to express emotions, studies on Speech Emotion Recognition (SER) have been conducted in various ways However, there are some areas for improvement in previous SER studies: (1) while some studies have performed multi-label classification, almost none have specifically utilized Korean speech data; (2) most studies have not utilized multiple features in combination for emotion recognition. Therefore, this study proposes deep fusion models for multi-label emotion classification using Korean speech data and follows four steps: (1) preprocessing speech data labeled with Sadness, Happiness, Neutral, Anger, and Disgust; (2) applying data augmentation to address the data imbalance and extracting speech features, including the Log-mel spectrogram, Mel-Frequency Cepstral Coefficients (MFCCs), and Voice Quality Features; (3) constructing models using deep fusion architectures; and (4) validating the performance of the constructed models. The experimental results demonstrated that the proposed model, which utilizes the Log-mel spectrogram and MFCCs with a fusion of Vision-Transformer and 1D Convolutional Neural Network–Long Short-Term Memory, achieved the highest average binary accuracy of 71.2% for multi-label classification, outperforming other baseline models. Consequently, this study anticipates that the proposed model will find application based on Korean speech, specifically mental healthcare and smart service systems.

Funder

Konkuk University

Publisher

MDPI AG

Link

https://www.mdpi.com/2076-3417/14/17/7604/pdf

Reference52 articles.

1. Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers;Speech Commun.,2020

2. Automatic, dimensional and continuous emotion recognition;Gunes;Int. J. Synth. Emot. (IJSE),2010

3. Artificial intelligence assisted improved human-computer interactions for computer systems;Alkatheiri;Comput. Electr. Eng.,2022

4. Jo, A.-H., and Kwak, K.-C. (2023). Speech emotion recognition based on two-stream deep learning model using Korean audio information. Appl. Sci., 13.

5. Ali, M., Mosa, A.H., Machot, F.A., and Kyamakya, K. (2018). Emotion recognition involving physiological and speech signals: A comprehensive review. Recent Advances in Nonlinear Dynamics and Synchronization: With Selected Applications in Electrical Engineering Neurocomputing, and Transportation, Springer.