Affiliation:
1. School of Computer Science, Civil Aviation Flight University of China, Guanghan 618307, China
Abstract
The quality of air traffic control speech is crucial. However, internal and external noise can impact air traffic control speech quality. Clear speech instructions and feedback help optimize flight processes and responses to emergencies. The traditional speech enhancement method based on a deep neural network and ideal ratio mask (DNN-IRM) is prone to distortion of the target speech in a strong noise environment. This paper introduces an air traffic control speech enhancement method based on an improved DNN-IRM. It employs LeakyReLU as an activation function to alleviate the gradient vanishing problem, improves the DNN network structure to enhance the IRM estimation capability, and adjusts the IRM weights to reduce noise interference in the target speech. The experimental results show that, compared with other methods, this method improves the perceptual evaluation of speech quality (PESQ), short-term objective intelligibility (STOI), scale-invariant signal-to-noise ratio (SI-SNR), and speech spectrogram clarity. In addition, we use this method to enhance real air traffic control speech, and the speech quality is also improved.
Funder
National Key R&D Program of China
Fundamental Research Funds for the Central Universities
Reference25 articles.
1. Peng, Y., Wen, X., Kong, J., Meng, Y., and Wu, M. (2023). A Study on the Normalized Delineation of Airspace Sectors Based on Flight Conflict Dynamics. Appl. Sci., 13.
2. Wu, Y., Li, G., and Fu, Q. (2023). Non-Intrusive Air Traffic Control Speech Quality Assessment with ResNet-BiLSTM. Appl. Sci., 13.
3. Identifying and managing risks of AI-driven operations: A case study of automatic speech recognition for improving air traffic safety;Yi;Chin. J. Aeronaut.,2023
4. Suppression of Acoustic Noise in Speech Using Spectral Subtraction;Boll;IEEE Trans. Acoust. Speech Signal Process.,1979
5. A Signal Subspace Approach for Speech Enhancement;Ephraim;IEEE Trans. Speech Audio Process.,1995