“In-Network Ensemble”: Deep Ensemble Learning with Diversified Knowledge Distillation-Reference-Cited by-同舟云学术

“In-Network Ensemble”: Deep Ensemble Learning with Diversified Knowledge Distillation

Published:2021-10-31 Issue:5 Volume:12 Page:1-19
ISSN:2157-6904
Container-title:ACM Transactions on Intelligent Systems and Technology
language:en
Short-container-title:ACM Trans. Intell. Syst. Technol.

Author:

Li Xingjian¹,Xiong Haoyi²,Chen Zeyu³,Huan Jun²,Xu Cheng-Zhong⁴,Dou Dejing²

Affiliation:

1. Baidu Inc., Beijing, China and University of Macau, Taipa, Macau, China

2. Baidu Inc., Beijing, China

3. Baidu Inc., Beijing, Guandong, China

4. University of Macau, Taipa, Macau, China

Abstract

Ensemble learning is a widely used technique to train deep convolutional neural networks (CNNs) for improved robustness and accuracy. While existing algorithms usually first train multiple diversified networks and then assemble these networks as an aggregated classifier, we propose a novel learning paradigm, namely, “In-Network Ensemble” ( INE ) that incorporates the diversity of multiple models through training a SINGLE deep neural network. Specifically, INE segments the outputs of the CNN into multiple independent classifiers, where each classifier is further fine-tuned with better accuracy through a so-called diversified knowledge distillation process . We then aggregate the fine-tuned independent classifiers using an Averaging-and-Softmax operator to obtain the final ensemble classifier. Note that, in the supervised learning settings, INE starts the CNN training from random, while, under the transfer learning settings, it also could start with a pre-trained model to incorporate the knowledge learned from additional datasets. Extensive experiments have been done using eight large-scale real-world datasets, including CIFAR, ImageNet, and Stanford Cars, among others, as well as common deep network architectures such as VGG, ResNet, and Wide ResNet. We have evaluated the method under two tasks: supervised learning and transfer learning. The results show that INE outperforms the state-of-the-art algorithms for deep ensemble learning with improved accuracy.

Funder

National Key Research and Development Program of China

Science and Technology Development Fund of Macau SAR

GuangDong Basic and Applied Basic Research Foundation

Key-area Research and Development Program of Guangdong Province

Publisher

Association for Computing Machinery (ACM)

Subject

Artificial Intelligence,Theoretical Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3473464

Reference51 articles.

1. Consistency of random forests and other averaging classifiers;Biau Gérard;J. Mach. Learn. Res. 9,2008

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. SeDPGK: Semi-supervised software defect prediction with graph representation learning and knowledge distillation;Information and Software Technology;2024-10

2. Guidelines for the Regularization of Gammas in Batch Normalization for Deep Residual Networks;ACM Transactions on Intelligent Systems and Technology;2024-03-29

3. Medical Image Classifications Using Convolutional Neural Networks: A Survey of Current Methods and Statistical Modeling of the Literature;Machine Learning and Knowledge Extraction;2024-03-21

4. PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models;2023 IEEE High Performance Extreme Computing Conference (HPEC);2023-09-25

5. Sequential or jumping: context-adaptive response generation for open-domain dialogue systems;Applied Intelligence;2022-09-02