Prompt guidance query with cascaded constraint decoders for human–object interaction detection-Reference-Cited by-同舟云学术

Prompt guidance query with cascaded constraint decoders for human–object interaction detection

Published:2024-03-29 Issue: Volume: Page:
ISSN:1751-9632
Container-title:IET Computer Vision
language:en
Short-container-title:IET Computer Vision

Author:

Liu Sheng¹,Guo Bingnan¹^ORCID,Zhang Feng¹,Chen Junhao¹,Chen Ruixiang¹

Affiliation:

1. College of Computer Science and Technology Zhejiang University of Technology Hangzhou China

Abstract

AbstractHuman–object interaction (HOI) detection, which localises and recognises interactions between human and object, requires high‐level image and scene understanding. Recent methods for HOI detection typically utilise transformer‐based architecture to build unified future representation. However, these methods use random initial queries to predict interactive human–object pairs, leading to a lack of prior knowledge. Furthermore, most methods provide unified features to forecast interactions using conventional decoder structures, but they lack the ability to build efficient multi‐task representations. To address these problems, we propose a novel two‐stage HOI detector called PGCD, mainly consisting of prompt guidance query and cascaded constraint decoders. Firstly, the authors propose a novel prompt guidance query generation module (PGQ) to introduce the guidance‐semantic features. In PGQ, the authors build visual‐semantic transfer to obtain fuller semantic representations. In addition, a cascaded constraint decoder architecture (CD) with random masks is designed to build fine‐grained interaction features and improve the model's generalisation performance. Experimental results demonstrate that the authors’ proposed approach obtains significant performance on the two widely used benchmarks, that is, HICO‐DET and V‐COCO.

Funder

Natural Science Foundation of Zhejiang Province

Publisher

Institution of Engineering and Technology (IET)

Reference45 articles.

1. Learning to Detect Human-Object Interactions

2. Gao C. Zou Y. Huang J.B.:ican: instance‐centric attention network for human‐object interaction detection.arXiv preprint arXiv:1808.10437. (2018)

3. Detecting and Recognizing Human-Object Interactions