Rethinking Centered Kernel Alignment in Knowledge Distillation

Zikai Zhou; Yunhang Shen; Shitong Shao; Linrui Gong; Shaohui Lin

arXiv:2401.11824·cs.CV·May 1, 2024·1 cites

Rethinking Centered Kernel Alignment in Knowledge Distillation

Zikai Zhou, Yunhang Shen, Shitong Shao, Linrui Gong, Shaohui Lin

PDF

Open Access 1 Repo

TL;DR

This paper reexamines the theoretical foundations of Centered Kernel Alignment (CKA) in knowledge distillation, proposing a new framework that simplifies and enhances its application for better performance in model compression tasks.

Contribution

It provides a theoretical analysis of CKA, introduces the RCKA framework linking CKA with MMD, and develops a task-adaptive method that improves distillation efficiency and effectiveness.

Findings

01

Achieves state-of-the-art results on CIFAR-100, ImageNet-1k, and MS-COCO.

02

Demonstrates that RCKA reduces computational costs while maintaining performance.

03

Validates the theoretical connection between CKA and MMD in knowledge distillation.

Abstract

Knowledge distillation has emerged as a highly effective method for bridging the representation discrepancy between large-scale models and lightweight models. Prevalent approaches involve leveraging appropriate metrics to minimize the divergence or distance between the knowledge extracted from the teacher model and the knowledge learned by the student model. Centered Kernel Alignment (CKA) is widely used to measure representation similarity and has been applied in several knowledge distillation methods. However, these methods are complex and fail to uncover the essence of CKA, thus not answering the question of how to use CKA to achieve simple and effective distillation properly. This paper first provides a theoretical perspective to illustrate the effectiveness of CKA, which decouples CKA to the upper bound of Maximum Mean Discrepancy~(MMD) and a constant term. Drawing from this, we…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

klayand/pcka
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Neural Network Applications · Domain Adaptation and Few-Shot Learning · Multimodal Machine Learning Applications

MethodsKnowledge Distillation