SOTAVerified

Knowledge Distillation

Knowledge distillation is the process of transferring knowledge from a large model to a smaller one. While large models (such as very deep neural networks or ensembles of many models) have higher knowledge capacity than small models, this capacity might not be fully utilized.

Papers

Showing 101–150 of 4240 papers

TitleStatusHype
Towards Low-Latency Event Stream-based Visual Object Tracking: A Slow-Fast ApproachCode0
Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image CorruptionCode0
LAMeTA: Intent-Aware Agentic Network Optimization via a Large AI Model-Empowered Two-Stage Approach—0
Always Clear Depth: Robust Monocular Depth Estimation under Adverse WeatherCode1
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning—0
On Membership Inference Attacks in Knowledge DistillationCode0
Denoising Mutual Knowledge Distillation in Bi-Directional Multiple Instance Learning—0
FiGKD: Fine-Grained Knowledge Distillation via High-Frequency Detail Transfer—0
Semantically-Aware Game Image Quality Assessment—0
Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors—0
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge DistillationCode0
Advancing Multiple Instance Learning with Continual Learning for Whole Slide Imaging—0
DCSNet: A Lightweight Knowledge Distillation-Based Model with Explainable AI for Lung Cancer Diagnosis from Histopathological Images—0
MoKD: Multi-Task Optimization for Knowledge Distillation—0
Low-Complexity Inference in Continual Learning via Compressed Knowledge Transfer—0
Foundation Models Knowledge Distillation For Battery Capacity Degradation ForecastCode1
Fusing Bidirectional Chains of Thought and Reward Mechanisms A Method for Enhancing Question-Answering Capabilities of Large Language Models for Chinese Intangible Cultural Heritage—0
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits—0
Topology-Guided Knowledge Distillation for Efficient Point Cloud ProcessingCode0
Channel Fingerprint Construction for Massive MIMO: A Deep Conditional Generative Approach—0
KDH-MLTC: Knowledge Distillation for Healthcare Multi-Label Text Classification—0
Ranking-aware Continual Learning for LiDAR Place Recognition—0
Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMsCode2
Simple Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head OptimizationCode0
Knowledge Distillation for Enhancing Walmart E-commerce Search Relevance Using Large Language Models—0
Human in the Latent Loop (HILL): Interactively Guiding Model Training Through Human Intuition—0
Robust & Precise Knowledge Distillation-based Novel Context-Aware Predictor for Disease Detection in Brain and Gastrointestinal—0
Federated Deconfounding and Debiasing Learning for Out-of-Distribution Generalization—0
Biomed-DPT: Dual Modality Prompt Tuning for Biomedical Vision-Language ModelsCode0
ABKD: Pursuing a Proper Allocation of the Probability Mass in Knowledge Distillation via α-β-DivergenceCode1
Theoretical Guarantees for LT-TTD: A Unified Transformer-based Architecture for Two-Level Ranking Systems—0
Action Spotting and Precise Event Detection in Sports: Datasets, Methods, and Challenges—0
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation—0
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance—0
Image Recognition with Online Lightweight Vision Transformer: A SurveyCode0
Artificial Behavior Intelligence: Technology, Challenges, and Future Directions—0
End-to-end fully-binarized network design: from Generic Learned Thermometer to Block Pruning—0
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks—0
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques—0
FedSDAF: Leveraging Source Domain Awareness for Enhanced Federated Domain GeneralizationCode0
Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge DistillationCode2
Segment Any RGB-Thermal Model with Language-aided Distillation—0
High-Fidelity Pseudo-label Generation by Large Language Models for Training Robust Radiology Report Classifiers—0
Toward Data-centric Directed Graph Learning: An Entropy-driven Approach—0
Llama-Nemotron: Efficient Reasoning Models—0
Uncertainty-Aware Multi-Expert Knowledge Distillation for Imbalanced Disease Grading—0
Enhancing New-item Fairness in Dynamic Recommender SystemsCode0
CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation—0
How to Backdoor the Knowledge Distillation—0
Head-Tail-Aware KL Divergence in Knowledge Distillation for Spiking Neural Networks—0
Show:102550
← PrevPage 3 of 85Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ScaleKD (T:BEiT-L S:ViT-B/14)Top-1 accuracy %86.43—Unverified
2ScaleKD (T:Swin-L S:ViT-B/16)Top-1 accuracy %85.53—Unverified
3ScaleKD (T:Swin-L S:ViT-S/16)Top-1 accuracy %83.93—Unverified
4ScaleKD (T:Swin-L S:Swin-T)Top-1 accuracy %83.8—Unverified
5KD++(T: regnety-16GF S:ViT-B)Top-1 accuracy %83.6—Unverified
6VkD (T:RegNety 160 S:DeiT-S)Top-1 accuracy %82.9—Unverified
7SpectralKD (T:Swin-S S:Swin-T)Top-1 accuracy %82.7—Unverified
8ScaleKD (T:Swin-L S:ResNet-50)Top-1 accuracy %82.55—Unverified
9DiffKD (T:Swin-L S: Swin-T)Top-1 accuracy %82.5—Unverified
10DIST (T: Swin-L S: Swin-T)Top-1 accuracy %82.3—Unverified
#ModelMetricClaimedVerifiedStatus
1SRD (T:resnet-32x4, S:shufflenet-v2)Top-1 Accuracy (%)79.86—Unverified
2shufflenet-v2(T:resnet-32x4, S:shufflenet-v2)Top-1 Accuracy (%)78.76—Unverified
3MV-MR (T: CLIP/ViT-B-16 S: resnet50)Top-1 Accuracy (%)78.6—Unverified
4resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)78.28—Unverified
5resnet8x4 (T: resnet32x4 S: resnet8x4 [modified])Top-1 Accuracy (%)78.08—Unverified
6ReviewKD++(T:resnet-32x4, S:shufflenet-v2)Top-1 Accuracy (%)77.93—Unverified
7ReviewKD++(T:resnet-32x4, S:shufflenet-v1)Top-1 Accuracy (%)77.68—Unverified
8resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)77.5—Unverified
9resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)76.68—Unverified
10resnet8x4 (T: resnet32x4 S: resnet8x4)Top-1 Accuracy (%)76.31—Unverified
#ModelMetricClaimedVerifiedStatus
1LSHFM (T: ResNet101 S: ResNet50)mAP93.17—Unverified
2LSHFM (T: ResNet101 S: MobileNetV2)mAP90.14—Unverified
#ModelMetricClaimedVerifiedStatus
1TIE-KD (T: Adabins S: MobileNetV2)RMSE2.43—Unverified