SOTAVerified

Knowledge Distillation

Knowledge distillation is the process of transferring knowledge from a large model to a smaller one. While large models (such as very deep neural networks or ensembles of many models) have higher knowledge capacity than small models, this capacity might not be fully utilized.

Papers

Showing 151–175 of 4240 papers

TitleStatusHype
Compressing Deep Graph Neural Networks via Adversarial Knowledge DistillationCode1
Computation-Efficient Knowledge Distillation via Uncertainty-Aware MixupCode1
Complementary Relation Contrastive DistillationCode1
Comprehensive Knowledge Distillation with Causal InterventionCode1
ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian SplattingCode1
Consensual Collaborative Training And Knowledge Distillation Based Facial Expression Recognition Under Noisy AnnotationsCode1
Contrastive Deep SupervisionCode1
Collaborative Distillation for Ultra-Resolution Universal Style TransferCode1
COMEDIAN: Self-Supervised Learning and Knowledge Distillation for Action Spotting using TransformersCode1
Model LEGO: Creating Models Like Disassembling and Assembling Building BlocksCode1
Coaching a Teachable StudentCode1
Communication-Efficient Federated Learning through Adaptive Weight Clustering and Server-Side DistillationCode1
CLRKDNet: Speeding up Lane Detection with Knowledge DistillationCode1
Cloud Object Detector Adaptation by Integrating Different Source KnowledgeCode1
CMDFusion: Bidirectional Fusion Network with Cross-modality Knowledge Distillation for LIDAR Semantic SegmentationCode1
CL-LoRA: Continual Low-Rank Adaptation for Rehearsal-Free Class-Incremental LearningCode1
CLIP model is an Efficient Continual LearnerCode1
FocusNet: Classifying Better by Focusing on Confusing ClassesCode1
CMD: Self-supervised 3D Action Representation Learning with Cross-modal Mutual DistillationCode1
Class-relation Knowledge Distillation for Novel Class DiscoveryCode1
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as TeachersCode1
Class-Incremental Learning by Knowledge Distillation with Adaptive Feature ConsolidationCode1
Class-Balanced Distillation for Long-Tailed Visual RecognitionCode1
Class-incremental Novel Class DiscoveryCode1
CLIP-guided Federated Learning on Heterogeneous and Long-Tailed DataCode1
Show:102550
← PrevPage 7 of 170Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ScaleKD (T:BEiT-L S:ViT-B/14)Top-1 accuracy %86.43—Unverified
2ScaleKD (T:Swin-L S:ViT-B/16)Top-1 accuracy %85.53—Unverified
3ScaleKD (T:Swin-L S:ViT-S/16)Top-1 accuracy %83.93—Unverified
4ScaleKD (T:Swin-L S:Swin-T)Top-1 accuracy %83.8—Unverified
5KD++(T: regnety-16GF S:ViT-B)Top-1 accuracy %83.6—Unverified
6VkD (T:RegNety 160 S:DeiT-S)Top-1 accuracy %82.9—Unverified
7SpectralKD (T:Swin-S S:Swin-T)Top-1 accuracy %82.7—Unverified
8ScaleKD (T:Swin-L S:ResNet-50)Top-1 accuracy %82.55—Unverified
9DiffKD (T:Swin-L S: Swin-T)Top-1 accuracy %82.5—Unverified
10DIST (T: Swin-L S: Swin-T)Top-1 accuracy %82.3—Unverified
#ModelMetricClaimedVerifiedStatus
1SRD (T:resnet-32x4, S:shufflenet-v2)Top-1 Accuracy (%)79.86—Unverified
2shufflenet-v2(T:resnet-32x4, S:shufflenet-v2)Top-1 Accuracy (%)78.76—Unverified
3MV-MR (T: CLIP/ViT-B-16 S: resnet50)Top-1 Accuracy (%)78.6—Unverified
4resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)78.28—Unverified
5resnet8x4 (T: resnet32x4 S: resnet8x4 [modified])Top-1 Accuracy (%)78.08—Unverified
6ReviewKD++(T:resnet-32x4, S:shufflenet-v2)Top-1 Accuracy (%)77.93—Unverified
7ReviewKD++(T:resnet-32x4, S:shufflenet-v1)Top-1 Accuracy (%)77.68—Unverified
8resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)77.5—Unverified
9resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)76.68—Unverified
10resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)76.31—Unverified
#ModelMetricClaimedVerifiedStatus
1LSHFM (T: ResNet101 S: ResNet50)mAP93.17—Unverified
2LSHFM (T: ResNet101 S: MobileNetV2)mAP90.14—Unverified
#ModelMetricClaimedVerifiedStatus
1TIE-KD (T: Adabins S: MobileNetV2)RMSE2.43—Unverified