Emotion Modeling in Speech Signals: Discrete Wavelet Transform and Machine Learning Tools for Emotion Recognition System-Reference-Cited by-同舟云学术

Emotion Modeling in Speech Signals: Discrete Wavelet Transform and Machine Learning Tools for Emotion Recognition System

Published:2024-01 Issue:1 Volume:2024 Page:
ISSN:1687-9724
Container-title:Applied Computational Intelligence and Soft Computing
language:en
Short-container-title:Applied Computational Intelligence and Soft Computing

Author:

Daqrouq K.^ORCID,Balamesh A.,Alrusaini O.,Alkhateeb A.,Balamash A. S.

Abstract

Speech emotion recognition (SER) is a challenging task due to the complex and subtle nature of emotions. This study proposes a novel approach for emotion modeling using speech signals by combining discrete wavelet transform (DWT) with linear prediction coding (LPC). The performance of various classifiers, including support vector machine (SVM), K‐Nearest Neighbors (KNN), Efficient Logistic Regression, Naive Bayes, Ensemble, and Neural Network, was evaluated for emotion classification using the EMO‐DB dataset. Evaluation metrics such as area under the curve (AUC), average prediction accuracy, and cross‐validation techniques were employed. The results indicate that KNN and SVM classifiers exhibited high accuracy in distinguishing sadness from other emotions. Ensemble methods and Neural Networks also demonstrated strong performance in sadness classification. While Efficient Logistic Regression and Naive Bayes classifiers showed competitive performance, they were slightly less accurate compared to other classifiers. Furthermore, the proposed feature extraction method yielded the highest average accuracy, and its combination with formants or wavelet entropy further improved classification accuracy. On the other hand, Efficient Logistic Regression exhibited the lowest accuracies among the classifiers. The uniqueness of this study was that it investigated a combined feature extraction method and integrated them to compare with various forms of combinations. However, the purposes of the investigation include improved performance of the classifiers, high effectiveness of the system, and the potential for emotion classification tasks. These findings can guide the selection of appropriate classifiers and feature extraction methods in future research and real‐world applications. Further investigations can focus on refining classifiers and exploring additional feature extraction techniques to enhance emotion classification accuracy.

Funder

King Abdulaziz University

Publisher

Wiley

Link

http://downloads.hindawi.com/journals/acisc/2024/7184018.pdf

Reference25 articles.

1. A Speech Emotion Recognition Model Based on Multi-Level Local Binary and Local Ternary Patterns

2. Speech Emotion Detection Through Live Calls

3. LalithaS. MudupuA. NandyalaB. V. andMunagalaR. Speech emotion recognition using DWT Proceedings of the 2015 IEEE International Conference on Computational Intelligence and Computing Research (ICCIC) December 2015 Madurai India IEEE 1–4 https://doi.org/10.1109/ICCIC.2015.7435630 2-s2.0-84965069900.

4. SasteS. T.andJagdaleS. M. Emotion recognition from speech using MFCC and DWT for security system Proceedings of the 2017 International conference of Electronics Communication and Aerospace Technology (ICECA) April 2017 Coimbatore India IEEE 701–704 https://doi.org/10.1109/ICECA.2017.8203631 2-s2.0-85042799558.

5. RamC. S.andPonnusamyR. An effective automatic speech emotion recognition for Tamil language based on DWT and MFCC using Stability-plasticity dilemma Neural network Proceedings of the International Conference on Information Communication and Embedded Systems (ICICES2014) February 2014 Chennai India IEEE 1–6 https://doi.org/10.1109/ICICES.2014.7034102 2-s2.0-84925339192.