SER: Performance Evaluation of CNN Model Along with an Overview of Available Indic Speech Datasets, and Transition of Classifiers From Traditional to Modern Era-Reference-Cited by-同舟云学术

SER: Performance Evaluation of CNN Model Along with an Overview of Available Indic Speech Datasets, and Transition of Classifiers From Traditional to Modern Era

Published:2023-06-26 Issue: Volume: Page:
ISSN:2375-4699
Container-title:ACM Transactions on Asian and Low-Resource Language Information Processing
language:en
Short-container-title:ACM Trans. Asian Low-Resour. Lang. Inf. Process.

Author:

Khurana Surbhi¹^ORCID,Dev Amita¹^ORCID,Bansal Poonam¹^ORCID

Affiliation:

1. Department of IT, Indira Gandhi Delhi Technical University for Women (IGDTUW), Delhi, India

Abstract

Speech emotion recognition (SER) is a rapidly evolving field in affective computing and human-computer interaction. In general, a SER system extracts and classifies prominent elements called features from a pre-processed speech signal to target the presence of speaker's certain emotion. This paper explores the utilization of deep learning classifiers in SER and surveys available datasets in both Indic and international languages. The paper highlights the significance of SER in enhancing human-computer interaction and presents deep learning as an effective approach to handle the complexity of speech signals. Various deep learning architectures, including Convolution Neural Networks (CNNs), Recurrent Neural Network (RNNs), and hybrid models, are analysed in terms of training methodology, and performance on benchmark datasets. Additionally, the paper conducts a comprehensive survey of publicly available datasets for speech emotion recognition, considering emotional categories, language diversity, recording conditions, and sample sizes. Challenges in adapting deep learning models to these datasets, such as data augmentation and cross-lingual transfer learning, are discussed. Moreover, the CNN based model is analysed on accuracy, precision, recall and F-1 score on Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) dataset with the value 84%, 85%, 84% and 84% resp. The review concludes with key findings, emphasizing the strengths and limitations of deep learning classifiers for SER. It identifies the need for standardized evaluation protocols, exploration of transfer learning across languages, and development of robust and culturally diverse datasets as future research directions.

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3605778

Reference87 articles.

1. Vocal Pitch during Simulated Emotion;F.;Lancet,1943

2. Emotion Communication System

3. Towards improving feature extraction and classification for activity recognition on streaming data

4. Speech Emotion Recognition Using Deep Learning Techniques: A Review

5. Feature learning via deep belief network for Chinese speech emotion recognition;Zhang S.;Commun. Comput. Inf. Sci.,2016

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ADAM optimised human speech emotion recogniser based on statistical information distribution of chroma, MFCC, and MBSE features;Multimedia Tools and Applications;2024-05-13