Frequency, Time, Representation and Modeling Aspects for Major Speech and Audio Processing Applications-Reference-Cited by-同舟云学术

Frequency, Time, Representation and Modeling Aspects for Major Speech and Audio Processing Applications

Published:2022-08-22 Issue:16 Volume:22 Page:6304
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Kacur Juraj^ORCID,Puterka Boris,Pavlovicova Jarmila^ORCID,Oravec Milos

Abstract

There are many speech and audio processing applications and their number is growing. They may cover a wide range of tasks, each having different requirements on the processed speech or audio signals and, therefore, indirectly, on the audio sensors as well. This article reports on tests and evaluation of the effect of basic physical properties of speech and audio signals on the recognition accuracy of major speech/audio processing applications, i.e., speech recognition, speaker recognition, speech emotion recognition, and audio event recognition. A particular focus is on frequency ranges, time intervals, a precision of representation (quantization), and complexities of models suitable for each class of applications. Using domain-specific datasets, eligible feature extraction methods and complex neural network models, it was possible to test and evaluate the effect of basic speech and audio signal properties on the achieved accuracies for each group of applications. The tests confirmed that the basic parameters do affect the overall performance and, moreover, this effect is domain-dependent. Therefore, accurate knowledge of the extent of these effects can be valuable for system designers when selecting appropriate hardware, sensors, architecture, and software for a particular application, especially in the case of limited resources.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/22/16/6304/pdf

Reference56 articles.

1. Automatic speech recognition: a survey

2. Two decades of speaker recognition evaluation at the national institute of standards and technology

3. A Real-Time Speech Emotion Recognition System and its Application in Online Learning

4. Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019

5. Fundamentals of Speech Recognition;Rabiner,1993

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Enabling Self-Powered Analog Voice Communication with Photovoltaic Cells and Optical Wireless Links Communication;2024 14th International Symposium on Communication Systems, Networks and Digital Signal Processing (CSNDSP);2024-07-17

2. Speaker age and gender recognition using 1D and 2D convolutional neural networks;Neural Computing and Applications;2023-11-28

3. Implementation of Real-Time Sound Source Localization using TMS320C6713 Board with Interaural Time Difference Method;2022 2nd International Seminar on Machine Learning, Optimization, and Data Science (ISMODE);2022-12-22