MLTU: Mixup Long-Tail Unsupervised Zero-Shot Image Classification on Vision-Language Models-Reference-Cited by-同舟云学术

MLTU: Mixup Long-Tail Unsupervised Zero-Shot Image Classification on Vision-Language Models

Published:2024-03-24 Issue: Volume: Page:
ISSN:
Container-title:
language:
Short-container-title:

Author:

Jia Yunpeng¹,Ye Xiufen¹,Mei Xinkui¹,Liu Yusong²,Guo Shuxiang³

Affiliation:

1. Harbin Engineering University

2. Harbin Medical University

3. Southern University of Science and Technology

Abstract

Vision-language models, such as Contrastive Language-Image Pretraining (CLIP), have demonstrated powerful capabilities in image classification under zero-shot settings. However, current Zero-Shot Learning (ZSL) relies on manually tagged samples of known classes through supervised learning, resulting in a waste of labor costs and limitations on foreseeable classes in real-world applications. To address these challenges, we propose the Mixup Long-Tail Unsupervised (MLTU) approach for open-world ZSL problems. The proposed approach employed a novel long-tail mixup loss that integrated class-based re-weighting assignments with a given mixup factor for each mixed visual embedding. To mitigate the adverse impact over time, we adopted a noisy learning strategy to filter out samples that generated incorrect labels. We reproduced the unsupervised results of existing state-of-the-art long-tail and noisy learning approaches. Experimental results demonstrate that MLTU achieves significant improvements in classification compared to these proven existing approaches on public datasets. Moreover, it serves as a plug-and-play solution for amending previous assignments and enhancing unsupervised performance. MLTU enables the automatic classification and correction of incorrect predictions caused by the projection bias of CLIP.

Publisher

Springer Science and Business Media LLC

Reference53 articles.

1. Larochelle, H., Erhan, D., Bengio, Y.: Zero-data learning of new tasks. In: AAAI, p. 3 (2008)

2. Attribute-based classification for zero-shot visual object categorization;Lampert CH;IEEE Transactions on Pattern Analysis and Machine Intelligence,2013

3. Li, K., Min, M., Fu, Y.: Rethinking zero-shot learning: A conditional visual classification perspective. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3583–3592 (2019)

4. Xu, W., Xian, Y., Wang, J., Schiele, B., Akata, Z.: Attribute prototype network for zero-shot learning. In: Advances in Neural Information Processing Systems, vol. 33, pp. 21969–21980 (2020)

5. Liang, J., Hu, D., Feng, J.: Domain adaptation with auxiliary target domain-oriented classifier. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16632–16642 (2021)