Noise-Robust Voice Conversion Using High-Quefrency Boosting via Sub-Band Cepstrum Conversion and Fusion-Reference-Cited by-同舟云学术

Noise-Robust Voice Conversion Using High-Quefrency Boosting via Sub-Band Cepstrum Conversion and Fusion

Published:2019-12-23 Issue:1 Volume:10 Page:151
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Miao Xiaokong^ORCID,Sun Meng^ORCID,Zhang Xiongwei,Wang Yimin

Abstract

This paper presents a noise-robust voice conversion method with high-quefrency boosting via sub-band cepstrum conversion and fusion based on the bidirectional long short-term memory (BLSTM) neural networks that can convert parameters of vocal tracks of a source speaker into those of a target speaker. With the implementation of state-of-the-art machine learning methods, voice conversion has achieved good performance given abundant clean training data. However, the quality and similarity of the converted voice are significantly degraded compared to that of a natural target voice due to various factors, such as limited training data and noisy input speech from the source speaker. To address the problem of noisy input speech, an architecture of voice conversion with statistical filtering and sub-band cepstrum conversion and fusion is introduced. The impact of noises on the converted voice is reduced by the accurate reconstruction of the sub-band cepstrum and the subsequent statistical filtering. By normalizing the mean and variance of the converted cepstrum to those of the target cepstrum in the training phase, a cepstrum filter was constructed to further improve the quality of the converted voice. The experimental results showed that the proposed method significantly improved the naturalness and similarity of the converted voice compared to the baselines, even with the noisy inputs of source speakers.

Funder

the National Natural Science Foundation of China

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/10/1/151/pdf

Reference32 articles.

1. Continuous probabilistic transform for voice conversion

2. Voice Conversion Based on Maximum-Likelihood Estimation of Spectral Parameter Trajectory

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. English Emotional Voice Conversion Using StarGAN Model;IEEE Access;2023

2. Noisy-to-Noisy Voice Conversion Under Variations of Noisy Condition;IEEE/ACM Transactions on Audio, Speech, and Language Processing;2023

3. Arabic Emotional Voice Conversion Using English Pre-Trained StarGANv2-VC-Based Model;Applied Sciences;2022-11-28

4. Speech Enhancement-assisted Voice Conversion in Noisy Environments;2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC);2022-11-07

5. Learning Noise-independent Speech Representation for High-quality Voice Conversion for Noisy Target Speakers;Interspeech 2022;2022-09-18