Reconstruction of Phonated Speech from Whispers Using Formant-Derived Plausible Pitch Modulation-Reference-Cited by-同舟云学术

Reconstruction of Phonated Speech from Whispers Using Formant-Derived Plausible Pitch Modulation

Published:2015-06-08 Issue:4 Volume:6 Page:1-21
ISSN:1936-7228
Container-title:ACM Transactions on Accessible Computing
language:en
Short-container-title:ACM Trans. Access. Comput.

Author:

Mcloughlin Ian V.¹,Sharifzadeh Hamid Reza²,Tan Su Lim³,Li Jingjie¹,Song Yan¹

Affiliation:

1. The University of Science and Technology of China, Hefei, Anhui, China

2. Unitec Institute of Technology, Auckland, New Zealand

3. Singapore Institute of Technology, Singapore

Abstract

Whispering is a natural, unphonated, secondary aspect of speech communications for most people. However, it is the primary mechanism of communications for some speakers who have impaired voice production mechanisms, such as partial laryngectomees, as well as for those prescribed voice rest, which often follows surgery or damage to the larynx. Unlike most people, who choose when to whisper and when not to, these speakers may have little choice but to rely on whispers for much of their daily vocal interaction. Even though most speakers will whisper at times, and some speakers can only whisper, the majority of today’s computational speech technology systems assume or require phonated speech. This article considers conversion of whispers into natural-sounding phonated speech as a noninvasive prosthetic aid for people with voice impairments who can only whisper. As a by-product, the technique is also useful for unimpaired speakers who choose to whisper. Speech reconstruction systems can be classified into those requiring training and those that do not. Among the latter, a recent parametric reconstruction framework is explored and then enhanced through a refined estimation of plausible pitch from weighted formant differences. The improved reconstruction framework, with proposed formant-derived artificial pitch modulation, is validated through subjective and objective comparison tests alongside state-of-the-art alternatives.

Funder

National Natural Science Foundation of China

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Science Applications,Human-Computer Interaction

Link

https://dl.acm.org/doi/pdf/10.1145/2737724

Reference31 articles.

1. Voice Rest after Microlaryngoscopy: Current Opinion and Practice

2. Speaker Recognition: Advancements and Challenges. Intech Book Publishers, Vienna, Austria;Beigi Homayoon;Chapter,2012

3. Hirose Hajime. 1986. Pathophysiology of motor speech disorders (dysarthria). Folia Phoniatrica et Logopaedica (International Journal of Phoniatrics Speech Therapy and Communication Pathology) 38 2--4 (June 1986) 61--88. Hirose Hajime. 1986. Pathophysiology of motor speech disorders (dysarthria). Folia Phoniatrica et Logopaedica (International Journal of Phoniatrics Speech Therapy and Communication Pathology) 38 2--4 (June 1986) 61--88.

4. Evaluation of Objective Quality Measures for Speech Enhancement

5. Reconstruction of whisper in Chinese by modified MELP

Cited by 14 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Automated Assessment of Glottal Dysfunction Through Unified Acoustic Voice Analysis;Journal of Voice;2020-09

2. Glottal Flow Synthesis for Whisper-to-Speech Conversion;IEEE/ACM Transactions on Audio, Speech, and Language Processing;2020

3. Effectiveness of Cross-Domain Architectures for Whisper-to-Normal Speech Conversion;2019 27th European Signal Processing Conference (EUSIPCO);2019-09

4. Whispered Speech to Normal Speech Conversion Using Bidirectional LSTMs with Meta-network;2019 IEEE 2nd International Conference on Information Communication and Signal Processing (ICICSP);2019-09

5. Whisper to Normal Speech Conversion Using Sequence-to-Sequence Mapping Model With Auditory Attention;IEEE Access;2019