Affiliation:
1. School of Computer Science and Technology Changchun University of Science and Technology Changchun China
2. Chongqing Research Institute Changchun University of Science and Technology Chongqing China
Abstract
AbstractDepth estimation from monocular panoramic image is a crucial step in 3D reconstruction, which is a close relationship with virtual reality and metaverse technologies. In recent years, some methods, such as HRDFuse, BiFuse++, and UniFuse, have employed a two‐branch neural network leveraging two common projections: equirectangular and cubemap projections (CMPs). The equirectangular projection (ERP) provides a complete field of view but introduces distortion, while the CMP avoids distortion but introduces discontinuity at the boundary of the cube. In order to address the issue of distortion and discontinuity, the authors propose an efficient depth estimation fusion module to balance the feature mapping of the two projections. Moreover, for the ERP, the authors propose a novel inflated network architecture to extend the receptive field and effectively harness visual information. Extensive experiments show that the authors’ method predicts more clear boundaries and accurate depth results while outperforming mainstream panoramic depth estimation algorithms.
Funder
Natural Science Foundation of Jilin Province
Publisher
Institution of Engineering and Technology (IET)
Subject
Electrical and Electronic Engineering,Computer Vision and Pattern Recognition,Signal Processing,Software