Peer Collaborative Learning for Online Knowledge Distillation-Reference-Cited by-同舟云学术

Peer Collaborative Learning for Online Knowledge Distillation

Published:2021-05-18 Issue:12 Volume:35 Page:10302-10310
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Wu Guile,Gong Shaogang

Abstract

Traditional knowledge distillation uses a two-stage training strategy to transfer knowledge from a high-capacity teacher model to a compact student model, which relies heavily on the pre-trained teacher. Recent online knowledge distillation alleviates this limitation by collaborative learning, mutual learning and online ensembling, following a one-stage end-to-end training fashion. However, collaborative learning and mutual learning fail to construct an online high-capacity teacher, whilst online ensembling ignores the collaboration among branches and its logit summation impedes the further optimisation of the ensemble teacher. In this work, we propose a novel Peer Collaborative Learning method for online knowledge distillation, which integrates online ensembling and network collaboration into a unified framework. Specifically, given a target network, we construct a multi-branch network for training, in which each branch is called a peer. We perform random augmentation multiple times on the inputs to peers and assemble feature representations outputted from peers with an additional classifier as the peer ensemble teacher. This helps to transfer knowledge from a high-capacity teacher to peers, and in turn further optimises the ensemble teacher. Meanwhile, we employ the temporal mean model of each peer as the peer mean teacher to collaboratively transfer knowledge among peers, which helps each peer to learn richer knowledge and facilitates to optimise a more stable model with better generalisation. Extensive experiments on CIFAR-10, CIFAR-100 and ImageNet show that the proposed method significantly improves the generalisation of various backbone networks and outperforms the state-of-the-art methods.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 42 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. DynamicKD: An effective knowledge distillation via dynamic entropy correction-based distillation for gap optimizing;Pattern Recognition;2024-09

2. KnowledgeIE: Unifying Online-Offline Distillation based on Knowledge Inheritance and Evolution;2024 International Joint Conference on Neural Networks (IJCNN);2024-06-30

3. Adversarial attacks and defenses for large language models (LLMs): methods, frameworks & challenges;International Journal of Multimedia Information Retrieval;2024-06-25

4. Learning From Human Educational Wisdom: A Student-Centered Knowledge Distillation Method;IEEE Transactions on Pattern Analysis and Machine Intelligence;2024-06

5. Federated Constrastive Learning and Visual Transformers for Personal Recommendation;Cognitive Computation;2024-05-08