PERFORMANCE ENHANCEMENT OF DEEP NEURAL NETWORK BASED AUTOMATIC VOICE DISORDER DETECTION SYSTEM WITH DATA AUGMENTATION — DETECTION OF LEUKOPLAKIA: A CASE STUDY

Author:

Thennal D. K.1,Nair Vrinda V.2ORCID,Indudharan R.3,Gopinath Deepa P.4ORCID

Affiliation:

1. Department of Computer Science, IIIT Kottayam, Kerala, India

2. State Project Facilitation Unit, Directorate of Technical Education, Trivandrum, Kerala, India

3. Consultant ENT Surgeon, Thrissur, Kerala, India

4. Department of Electronics and Communication Engineering, Government Engineering College Barton Hill, Trivandrum, Kerala, India

Abstract

Laryngeal pathologies resulting in voice disorders are normally diagnosed using invasive methods such as rigid laryngoscopy, flexible nasopharyngo-laryngoscopy and stroboscopy, which are expensive, time-consuming and often inconvenient to patients. Automatic Voice Disorder Detection (AVDD) systems are used for non-invasive screening to give an indicative direction to the physician as a preliminary diagnosis. Deep neural networks, known for their superior discrimination capabilities, can be used for AVDD Systems, provided there are sufficient samples for training. The most popular datasets used for developing AVDD systems lack sufficient samples in several pathological categories. Leukoplakia — a premalignant lesion, which may progress to carcinoma unless detected early — is one such pathology. Data augmentation is a technique used in deep learning environments to increase the size of the training datasets which lack sufficient samples for effective data analysis and classification. This study aims at investigating the performance enhancement of a deep learning-based AVDD system through a novel time domain data augmentation technique named ‘TempAug’. This method segments each data sample into short voice segments, so as to get multiple data from each sample, thereby generating a larger database (augmented database) for training a deep learning model. A deep neural network model, Long Short-Term Memory (LSTM) with Short Term Fourier Transform (STFT) coefficients as input features for classification, was used in this study for the detection of the voice disorder Leukoplakia. A series of experiments were done to investigate the effect of data augmentation and to find the optimum duration for segmentation. Based on experimental results, a detection strategy was developed and evaluated using an AVDD system, which gave an accuracy of 81.25%. The percentage increase in accuracy was found to be 46.9% with respect to the accuracy obtained for unaugmented data.

Publisher

National Taiwan University

Subject

Biomedical Engineering,Bioengineering,Biophysics

Cited by 1 articles. 订阅此论文施引文献 订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献

1. Deep Learning-based Test Data Augmentation Technology;2023 16th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI);2023-10-28

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3