Complementary Coarse-to-Fine Matching for Video Object Segmentation-Reference-Cited by-同舟云学术

Complementary Coarse-to-Fine Matching for Video Object Segmentation

Published:2023-07-12 Issue:6 Volume:19 Page:1-21
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Chen Zhen¹^ORCID,Yang Ming²^ORCID,Zhang Shiliang¹^ORCID

Affiliation:

1. National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University, China

2. Multimodality Cognition, Ant Group, USA

Abstract

Semi-supervised Video Object Segmentation (VOS) needs to establish pixel-level correspondences between a video frame and preceding segmented frames to leverage their segmentation clues. Most works rely on features at a single scale to establish those correspondences, e.g., perform dense matching with Convolutional Neural Network (CNN) features from a deep layer. Differently, this work explores complementary features at different scales to pursue more robust feature matching. A coarse feature from a deep layer is first adopted to get coarse pixel-level correspondences. We hence evaluate the quality of those correspondences, and select pixels with low-quality correspondences for fine-scale feature matching. Segmentation clues of previous frames are propagated by both coarse and fine-scale correspondences, which are fused with appearance features for object segmentation. Compared with previous works, this coarse-to-fine matching scheme is more robust to distractions by similar objects and better preserves object details. The sparse fine-scale matching also ensures a fast inference speed. On popular VOS datasets including DAVIS and YouTube-VOS, the proposed method shows promising performance compared with recent works.

Funder

Natural Science Foundation of China

The National Key Research and Development Program of China

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications,Hardware and Architecture

Link

https://dl.acm.org/doi/pdf/10.1145/3596496

Reference56 articles.

1. CNN in MRF: Video Object Segmentation via Inference in a CNN-Based Higher-Order Spatio-Temporal MRF

2. Goutam Bhat, Felix Järemo Lawin, Martin Danelljan, Andreas Robinson, Michael Felsberg, Luc Van Gool, and Radu Timofte. 2020. Learning what to learn for video object segmentation. In Computer VisionECCV 2020: 16th European Conference, Glasgow, UK, August 2328, 2020, Proceedings, Part II 16.

3. One-Shot Video Object Segmentation

4. Optimizing Video Object Detection via a Scale-Time Lattice

5. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Learning Nighttime Semantic Segmentation the Hard Way;ACM Transactions on Multimedia Computing, Communications, and Applications;2024-05-16