CLG: Contrastive Label Generation with Knowledge for Few-Shot Learning-Reference-Cited by-同舟云学术

CLG: Contrastive Label Generation with Knowledge for Few-Shot Learning

Published:2024-02-01 Issue:3 Volume:12 Page:472
ISSN:2227-7390
Container-title:Mathematics
language:en
Short-container-title:Mathematics

Author:

Ma Han¹^ORCID,Fan Baoyu¹^ORCID,Ng Benjamin K.¹,Lam Chan-Tong¹^ORCID

Affiliation:

1. Faculty of Applied Sciences, Macao Polytechnic University, Macao 999078, China

Abstract

Training large-scale models needs big data. However, the few-shot problem is difficult to resolve due to inadequate training data. It is valuable to use only a few training samples to perform the task, such as using big data for application scenarios due to cost and resource problems. So, to tackle this problem, we present a simple and efficient method, contrastive label generation with knowledge for few-shot learning (CLG). Specifically, we: (1) Propose contrastive label generation to align the label with data input and enhance feature representations; (2) Propose a label knowledge filter to avoid noise during injection of the explicit knowledge into the data and label; (3) Employ label logits mask to simplify the task; (4) Employ multi-task fusion loss to learn different perspectives from the training set. The experiments demonstrate that CLG achieves an accuracy of 59.237%, which is more than about 3% in comparison with the best baseline. It shows that CLG obtains better features and gives the model more information about the input sentences to improve the classification ability.

Funder

Macao Polytechnic University

Publisher

MDPI AG

Link

https://www.mdpi.com/2227-7390/12/3/472/pdf

Reference57 articles.

1. Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv.

2. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv.

3. Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019). Albert: A lite bert for self-supervised learning of language representations. arXiv.

4. Xlnet: Generalized autoregressive pretraining for language understanding;Yang;Adv. Neural Inf. Process. Syst.,2019

5. Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. (2019). Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv.