Encoder–Decoder Structure Fusing Depth Information for Outdoor Semantic Segmentation-Reference-Cited by-同舟云学术

Encoder–Decoder Structure Fusing Depth Information for Outdoor Semantic Segmentation

Published:2023-09-01 Issue:17 Volume:13 Page:9924
ISSN:2076-3417
Container-title:Applied Sciences
language:en
Short-container-title:Applied Sciences

Author:

Chen Songnan¹^ORCID,Tang Mengxia²,Dong Ruifang²,Kan Jiangming²

Affiliation:

1. School of Mathematics and Computer Science, Wuhan Polytechnic University, No.36 Huanhu Middle Road, Dongxihu District, Wuhan 430048, China

2. School of Technology, Beijing Forestry University, No.35 Qinghua East Road, Haidian District, Beijing 100083, China

Abstract

The semantic segmentation of outdoor images is the cornerstone of scene understanding and plays a crucial role in the autonomous navigation of robots. Although RGB–D images can provide additional depth information for improving the performance of semantic segmentation tasks, current state–of–the–art methods directly use ground truth depth maps for depth information fusion, which relies on highly developed and expensive depth sensors. Aiming to solve such a problem, we proposed a self–calibrated RGB-D image semantic segmentation neural network model based on an improved residual network without relying on depth sensors, which utilizes multi-modal information from depth maps predicted with depth estimation models and RGB image fusion for image semantic segmentation to enhance the understanding of a scene. First, we designed a novel convolution neural network (CNN) with an encoding and decoding structure as our semantic segmentation model. The encoder was constructed using IResNet to extract the semantic features of the RGB image and the predicted depth map and then effectively fuse them with the self–calibration fusion structure. The decoder restored the resolution of the output features with a series of successive upsampling structures. Second, we presented a feature pyramid attention mechanism to extract the fused information at multiple scales and obtain features with rich semantic information. The experimental results using the publicly available Cityscapes dataset and collected forest scene images show that our model trained with the estimated depth information can achieve comparable performance to the ground truth depth map in improving the accuracy of the semantic segmentation task and even outperforming some competitive methods.

Funder

National Natural Science Foundation of China

Science and Technology Fund of Henan Province

Research and Innovation Initiatives of WHPU

research funding from Wuhan Polytechnic University

Publisher

MDPI AG

Subject

Fluid Flow and Transfer Processes,Computer Science Applications,Process Chemistry and Technology,General Engineering,Instrumentation,General Materials Science

Link

https://www.mdpi.com/2076-3417/13/17/9924/pdf

Reference43 articles.

1. Xu, Y., Wang, H., Liu, X., He, H.R., Gu, Q., and Sun, W. (2019). Learning to See the Hidden Part of the Vehicle in the Autopilot Scene. Electronics, 8.

2. Scene terrain classification for autonomous vehicle navigation based on semantic segmentation method;Fusic;Trans. Inst. Meas. Control,2022

3. Explainable multi–module semantic guided attention based network for medical image segmentation;Karri;Comput. Biol. Med.,2022

4. CCTseg: A cascade composite transformer semantic segmentation network for UAV visual perception;Yi;Measurement,2022

5. A threshold selection method from gray–level histograms;Otsu;IEEE Trans. Syst. Man Cybern.,1979

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. CLGFormer: Cross-Level-Guided transformer for RGB-D semantic segmentation;Multimedia Tools and Applications;2024-05-09

2. CGAN-Based Forest Scene 3D Reconstruction from a Single Image;Forests;2024-01-18