Aligning Image Semantics and Label Concepts for Image Multi-Label Classification-Reference-Cited by-同舟云学术

Aligning Image Semantics and Label Concepts for Image Multi-Label Classification

Published:2023-02-06 Issue:2 Volume:19 Page:1-23
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Zhou Wei¹^ORCID,Xia Zhiwu¹^ORCID,Dou Peng¹^ORCID,Su Tao¹^ORCID,Hu Haifeng¹^ORCID

Affiliation:

1. School of Electronics and Information Technology, Sun Yat-sen University, Guangdong, People’s Republic of China

Abstract

Image multi-label classification task is mainly to correctly predict multiple object categories in the images. To capture the correlation between labels, graph convolution network based methods have to manually count the label co-occurrence probability from training data to construct a pre-defined graph as the input of graph network, which is inflexible and may degrade model generalizability. Moreover, most of the current methods cannot effectively align the learned salient object features with the label concepts, so that the predicted results of model may not be consistent with the image content. Therefore, how to learn the salient semantic features of images and capture the correlation between labels, and then effectively align them is one of the key to improve the performance of image multi-label classification task. To this end, we propose a novel image multi-label classification framework which aims to align I mage S emantics with L abel C oncepts ( ISLC ). Specifically, we propose a residual encoder to learn salient object features in the images, and exploit the self-attention layer in aligned decoder to automatically capture the correlation between labels. Then, we leverage the cross-attention layers in aligned decoder to align image semantic features with label concepts, so as to make the labels predicted by model more consistent with image content. Finally, the output features of the last layer of residual encoder and aligned decoder are fused to obtain the final output feature for classification. The proposed ISLC model achieves good performance on various prevalent multi-label image datasets such as MS-COCO 2014, PASCAL VOC 2007, VG-500, and NUS-WIDE with 87.2%, 96.9%, 39.4%, and 64.2%, respectively.

Funder

National Natural Science Foundation of China

Science and Technology Program of Guangdong Province

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications,Hardware and Architecture

Link

https://dl.acm.org/doi/pdf/10.1145/3550278

Reference74 articles.

1. Hakan Cevikalp, Burak Benligiray, Ömer Nezih Gerek, and Hasan Saribas. 2019. Semi-supervised robust deep neural networks for multi-label classification. In Proceedings of the CVPR Workshops. 9–17.

2. Knowledge-guided multi-label few-shot learning for general image recognition;Chen Tianshui;IEEE Transactions on Pattern Analysis and Machine Intelligence,2020

3. Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition

4. Multi-Label Image Recognition With Graph Convolutional Networks

5. Xiangxiang Chu Bo Zhang Zhi Tian Xiaolin Wei and Huaxia Xia. 2021. Do we really need explicit position encodings for vision transformers? arXiv:2102.10882. Retrieved from https://arxiv.org/abs/2102.10882.

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Multi-label recognition in open driving scenarios based on bipartite-driven superimposed dynamic graph;Image and Vision Computing;2024-09

2. Decoupling Deep Learning for Enhanced Image Recognition Interpretability;ACM Transactions on Multimedia Computing, Communications, and Applications;2024-07-10

3. Semantic deep learning and adaptive clustering for handling multimodal multimedia information retrieval;Multimedia Tools and Applications;2024-05-25

4. DATran: Dual Attention Transformer for Multi-Label Image Classification;IEEE Transactions on Circuits and Systems for Video Technology;2024-01

5. Causal multi-label learning for image classification;Neural Networks;2023-10