Recognition of Emotion with Intensity from Speech Signal Using 3D Transformed Feature and Deep Learning-Reference-Cited by-同舟云学术

Recognition of Emotion with Intensity from Speech Signal Using 3D Transformed Feature and Deep Learning

Published:2022-07-28 Issue:15 Volume:11 Page:2362
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Islam Md. Riadul^ORCID,Akhand M. A. H.^ORCID,Kamal Md Abdus Samad^ORCID,Yamada Kou

Abstract

Speech Emotion Recognition (SER), the extraction of emotional features with the appropriate classification from speech signals, has recently received attention for its emerging social applications. Emotional intensity (e.g., Normal, Strong) for a particular emotional expression (e.g., Sad, Angry) has a crucial influence on social activities. A person with intense sadness or anger may fall into severe disruptive action, eventually triggering a suicidal or devastating act. However, existing Deep Learning (DL)-based SER models only consider the categorization of emotion, ignoring the respective emotional intensity, despite its utmost importance. In this study, a novel scheme for Recognition of Emotion with Intensity from Speech (REIS) is developed using the DL model by integrating three speech signal transformation methods, namely Mel-frequency Cepstral Coefficient (MFCC), Short-time Fourier Transform (STFT), and Chroma STFT. The integrated 3D form of transformed features from three individual methods is fed into the DL model. Moreover, under the proposed REIS, both the single and cascaded frameworks with DL models are investigated. A DL model consists of a 3D Convolutional Neural Network (CNN), Time Distribution Flatten (TDF) layer, and Bidirectional Long Short-term Memory (Bi-LSTM) network. The 3D CNN block extracts convolved features from 3D transformed speech features. The convolved features were flattened through the TDF layer and fed into Bi-LSTM to classify emotion with intensity in a single DL framework. The 3D transformed feature is first classified into emotion categories in the cascaded DL framework using a DL model. Then, using a different DL model, the intensity level of the identified categories is determined. The proposed REIS has been evaluated on the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) benchmark dataset, and the cascaded DL framework is found to be better than the single DL framework. The proposed REIS method has shown remarkable recognition accuracy, outperforming related existing methods.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering

Link

https://www.mdpi.com/2079-9292/11/15/2362/pdf

Reference47 articles.

1. Design, analysis and experimental evaluation of block based transformation in MFCC computation for speaker recognition

2. The Feedforward Short-Time Fourier Transform

3. Hybrid Deep Network Scheme for Emotion Recognition in Speech

4. A CNN-Assisted Enhanced Audio Signal Processing for Speech Emotion Recognition

5. BanglaSER: A speech emotion recognition dataset for the Bangla language

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Feature Based Classifier for Bangla Currency Using Deep Learning;2024 6th International Conference on Electrical Engineering and Information & Communication Technology (ICEEICT);2024-05-02

2. Non-permissible Mobile Detection to Enhance Security in Bangladeshi Museum: A Multiprocess YOLO-Based Approach;2024 6th International Conference on Electrical Engineering and Information & Communication Technology (ICEEICT);2024-05-02

3. Integrating Large Language Models (LLMs) and Deep Representations of Emotional Features for the Recognition and Evaluation of Emotions in Spoken English;Applied Sciences;2024-04-23

4. Emotion recognition from EEG signal enhancing feature map using partial mutual information;Biomedical Signal Processing and Control;2024-02

5. KBES: A dataset for realistic Bangla speech emotion recognition with intensity level;Data in Brief;2023-12