Investigations on the Optimal Estimation of Speech Envelopes for the Two-Stage Speech Enhancement-Reference-Cited by-同舟云学术

Investigations on the Optimal Estimation of Speech Envelopes for the Two-Stage Speech Enhancement

Published:2023-07-16 Issue:14 Volume:23 Page:6438
ISSN:1424-8220
Container-title:Sensors
language:en
Short-container-title:Sensors

Author:

Song Yanjue¹^ORCID,Madhu Nilesh¹^ORCID

Affiliation:

1. IDLab, Ghent University—imec, 9000 Gent, Belgium

Abstract

Using the source-filter model of speech production, clean speech signals can be decomposed into an excitation component and an envelope component that is related to the phoneme being uttered. Therefore, restoring the envelope of degraded speech during speech enhancement can improve the intelligibility and quality of output. As the number of phonemes in spoken speech is limited, they can be adequately represented by a correspondingly limited number of envelopes. This can be exploited to improve the estimation of speech envelopes from a degraded signal in a data-driven manner. The improved envelopes are then used in a second stage to refine the final speech estimate. Envelopes are typically derived from the linear prediction coefficients (LPCs) or from the cepstral coefficients (CCs). The improved envelope is obtained either by mapping the degraded envelope onto pre-trained codebooks (classification approach) or by directly estimating it from the degraded envelope (regression approach). In this work, we first investigate the optimal features for envelope representation and codebook generation by a series of oracle tests. We demonstrate that CCs provide better envelope representation compared to using the LPCs. Further, we demonstrate that a unified speech codebook is advantageous compared to the typical codebook that manually splits speech and silence as separate entries. Next, we investigate low-complexity neural network architectures to map degraded envelopes to the optimal codebook entry in practical systems. We confirm that simple recurrent neural networks yield good performance with a low complexity and number of parameters. We also demonstrate that with a careful choice of the feature and architecture, a regression approach can further improve the performance at a lower computational cost. However, as also seen from the oracle tests, the benefit of the two-stage framework is now chiefly limited by the statistical noise floor estimate, leading to only a limited improvement in extremely adverse conditions. This highlights the need for further research on joint estimation of speech and noise for optimum enhancement.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Biochemistry,Instrumentation,Atomic and Molecular Physics, and Optics,Analytical Chemistry

Link

https://www.mdpi.com/1424-8220/23/14/6438/pdf

Reference28 articles.

1. Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator;Ephraim;IEEE Trans. Acoust. Speech Signal Process.,1984

2. Speech enhancement using a minimum mean-square error log-spectral amplitude estimator;Ephraim;IEEE Trans. Acoust. Speech Signal Process.,1985

3. Improved signal-to-noise ratio estimation for speech enhancement;Plapous;IEEE Trans. Audio Speech Lang. Process.,2006

4. Improved CEM for speech harmonic enhancement in single channel noise suppression;Song;IEEE/ACM Trans. Audio Speech Lang. Process.,2022

5. Instantaneous a priori SNR estimation by cepstral excitation manipulation;Elshamy;IEEE/ACM Trans. Audio Speech Lang. Process.,2017