Teacher-student complementary sample contrastive distillation.

Researchers

Journal

Neural networks : the official journal of the International Neural Network Society

Modalities

Models

Contrastive Sample Hardness Knowledge Distillation Student Self-Learning Supervision Signal Correction

Abstract

Knowledge distillation (KD) is a widely adopted model compression technique for improving the performance of compact student models, by utilizing the “dark knowledge” of a large teacher model. However, previous studies have not adequately investigated the effectiveness of supervision from the teacher model, and overconfident predictions in the student model may degrade its performance. In this work, we propose a novel framework, Teacher-Student Complementary Sample Contrastive Distillation (TSCSCD), that alleviate these challenges. TSCSCD consists of three key components: Contrastive Sample Hardness (CSH), Supervision Signal Correction (SSC), and Student Self-Learning (SSL). Specifically, CSH evaluates the teacher’s supervision for each sample by comparing the predictions of two compact models, one distilled from the teacher and the other trained from scratch. SSC corrects weak supervision according to CSH, while SSL employs integrated learning among multi-classifiers to regularize overconfident predictions. Extensive experiments on four real-world datasets demonstrate that TSCSCD outperforms recent state-of-the-art knowledge distillation techniques.Copyright © 2023 Elsevier Ltd. All rights reserved.

Show Full Text

Teacher-student complementary sample contrastive distillation.

Researchers

Journal

Modalities

Models

Abstract

Monitoring gamma type-I censored data using an exponentially weighted moving average control chart based on deep learning networks.

Coupled Nonlinear Delay Systems as Deep Convolutional Neural Networks.

A Combination Model of Shifting Joint Angle Changes With 3D-Deep Convolutional Neural Network to Recognize Human Activity.

Research on motion recognition based on multi-dimensional sensing data and deep learning algorithms.

Whole Process Monitoring Based on Unstable Neuron Output Information in Hidden Layers of Deep Belief Network.

Generalization analysis of deep CNNs under maximum correntropy criterion.

Leave a Reply Cancel reply

Researchers

Journal

Modalities

Models

Abstract

Similar Posts

Leave a Reply Cancel reply