SOTAVerified

Model Compression

Model Compression is an actively pursued area of research over the last few years with the goal of deploying state-of-the-art deep networks in low-power and resource limited devices without significant drop in accuracy. Parameter pruning, low-rank factorization and weight quantization are some of the proposed methods to compress the size of deep networks.

Source: KD-MRI: A knowledge distillation framework for image reconstruction and image restoration in MRI workflow

Papers

Showing 401450 of 1356 papers

TitleStatusHype
Foundations of Large Language Model Compression -- Part 1: Weight QuantizationCode0
Generalizing Teacher Networks for Effective Knowledge Distillation Across Student ArchitecturesCode0
Robust and Large-Payload DNN Watermarking via Fixed, Distribution-Optimized, WeightsCode0
Robust Knowledge Distillation Based on Feature Variance Against Backdoored Teacher ModelCode0
Robustness and Diversity Seeking Data-Free Knowledge DistillationCode0
Rotation Invariant Quantization for Model CompressionCode0
Few Shot Network Compression via Cross DistillationCode0
Adversarial Robustness vs. Model Compression, or Both?Code0
Lottery Aware Sparsity Hunting: Enabling Federated Learning on Resource-Limited EdgeCode0
What Do Compressed Deep Neural Networks Forget?Code0
Finding Deviated Behaviors of the Compressed DNN Models for Image ClassificationsCode0
Faithful Label-free Knowledge DistillationCode0
Model Compression with Adversarial Robustness: A Unified Optimization FrameworkCode0
Semi-Online Knowledge DistillationCode0
Compression-aware Continual Learning using Singular Value DecompositionCode0
Explicit-NeRF-QA: A Quality Assessment Database for Explicit NeRF Model CompressionCode0
Exact Backpropagation in Binary Weighted Networks with Group Weight TransformationsCode0
Exploiting Kernel Sparsity and Entropy for Interpretable CNN CompressionCode0
Compressing Vision Transformers for Low-Resource Visual LearningCode0
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech TransformersCode0
Enhancing In-Context Learning Performance with just SVD-Based Weight Pruning: A Theoretical PerspectiveCode0
Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and EfficiencyCode0
Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution AlignmentCode0
Exploring Gradient Flow Based Saliency for DNN Model CompressionCode0
EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware OptimizationCode0
Empirical Evaluation of Deep Learning Model Compression Techniques on the WaveNet VocoderCode0
ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer AccelerationCode0
An exploration of the effect of quantisation on energy consumption and inference time of StarCoder2Code0
Causal Explanation of Convolutional Neural NetworksCode0
StrassenNets: Deep Learning with a Multiplication BudgetCode0
Einconv: Exploring Unexplored Tensor Network Decompositions for Convolutional Neural NetworksCode0
Improved Knowledge Distillation via Full Kernel Matrix TransferCode0
Compressing Convolutional Neural Networks via Factorized Convolutional FiltersCode0
Efficient model compression with Random Operation Access Specific Tile (ROAST) hashingCode0
Efficient Speech Translation through Model Compression and Knowledge DistillationCode0
Systematic Outliers in Large Language ModelsCode0
Exploring Unexplored Tensor Network Decompositions for Convolutional Neural NetworksCode0
Tensorized Embedding Layers for Efficient Model CompressionCode0
Learning Accurate Performance Predictors for Ultrafast Automated Model CompressionCode0
On Model Compression for Neural Networks: Framework, Algorithm, and Convergence GuaranteeCode0
Compressed models are NOT miniature versions of large models0
Artemis: HE-Aware Training for Efficient Privacy-Preserving Machine Learning0
Comprehensive Survey of Model Compression and Speed up for Vision Transformers0
Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices0
Comprehensive Study on Performance Evaluation and Optimization of Model Compression: Bridging Traditional Deep Learning and Large Language Models0
Compositionality Unlocks Deep Interpretable Models0
A Comprehensive Review and a Taxonomy of Edge Machine Learning: Requirements, Paradigms, and Techniques0
Accelerating Very Deep Convolutional Networks for Classification and Detection0
CompMarkGS: Robust Watermarking for Compressed 3D Gaussian Splatting0
Complexity-Driven CNN Compression for Resource-constrained Edge AI0
Show:102550
← PrevPage 9 of 28Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MobileBERT + 2bit-1dim model compression using DKMAccuracy82.13Unverified
2MobileBERT + 1bit-1dim model compression using DKMAccuracy63.17Unverified