Personalized HRTF Modeling Based on Deep Neural Network Using Anthropometric Measurements and Images of the Ear-Reference-Cited by-同舟云学术

Personalized HRTF Modeling Based on Deep Neural Network Using Anthropometric Measurements and Images of the Ear

Published:2018-11-07 Issue:11 Volume:8 Page:2180
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Lee Geon,Kim Hong

Abstract

This paper proposes a personalized head-related transfer function (HRTF) estimation method based on deep neural networks by using anthropometric measurements and ear images. The proposed method consists of three sub-networks for representing personalized features and estimating the HRTF. As input features for neural networks, the anthropometric measurements regarding the head and torso are used for a feedforward deep neural network (DNN), and the ear images are used for a convolutional neural network (CNN). After that, the outputs of these two sub-networks are merged into another DNN for estimation of the personalized HRTF. To evaluate the performance of the proposed method, objective and subjective evaluations are conducted. For the objective evaluation, the root mean square error (RMSE) and the log spectral distance (LSD) between the reference HRTF and the estimated one are measured. Consequently, the proposed method provides the RMSE of −18.40 dB and LSD of 4.47 dB, which are lower by 0.02 dB and higher by 0.85 dB than the DNN-based method using anthropometric data without pinna measurements, respectively. Next, a sound localization test is performed for the subjective evaluation. As a result, it is shown that the proposed method can localize sound sources with higher accuracy of around 11% and 6% than the average HRTF method and DNN-based method, respectively. In addition, the reductions of the front/back confusion rate by 12.5% and 2.5% are achieved by the proposed method, compared to the average HRTF method and DNN-based method, respectively.

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

http://www.mdpi.com/2076-3417/8/11/2180/pdf

Reference38 articles.

1. Spatial Audio;Rumsey,2001

2. Spatial Hearing: The Psychophysics of Human Sound Localization;Blauert,1997

3. Factors That Influence the Localization of Sound in the Vertical Plane

4. Auditory distance perception in rooms

5. 3D Sound for Virtual Reality and Multimedia;Begault,1994

Cited by 40 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Auditory localization: a comprehensive practical review;Frontiers in Psychology;2024-07-10

2. Efficient prediction of individual head-related transfer functions based on 3D meshes;Applied Acoustics;2024-03

3. End-to-End Paired Ambisonic-Binaural Audio Rendering;IEEE/CAA Journal of Automatica Sinica;2024-02

4. HRTF Interpolation Using a Spherical Neural Process Meta-Learner;IEEE/ACM Transactions on Audio, Speech, and Language Processing;2024

5. Modeling of Individual Head-Related Transfer Functions (HRTFs) Based on Spatiotemporal and Anthropometric Features Using Deep Neural Networks;IEEE Access;2024