SOTAVerified

Model Compression

Model Compression is an actively pursued area of research over the last few years with the goal of deploying state-of-the-art deep networks in low-power and resource limited devices without significant drop in accuracy. Parameter pruning, low-rank factorization and weight quantization are some of the proposed methods to compress the size of deep networks.

Source: KD-MRI: A knowledge distillation framework for image reconstruction and image restoration in MRI workflow

Papers

Showing 301–350 of 1356 papers

TitleStatusHype
Blending LSTMs into CNNs—0
DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers—0
BioNetExplorer: Architecture-Space Exploration of Bio-Signal Processing Deep Neural Networks for Wearables—0
An Efficient Method of Training Small Models for Regression Problems with Knowledge Distillation—0
An Effective Information Theoretic Framework for Channel Pruning—0
BinaryBERT: Pushing the Limit of BERT Quantization—0
AdaKD: Dynamic Knowledge Distillation of ASR models using Adaptive Loss Weighting—0
Distilling Optimal Neural Networks: Rapid Search in Diverse Spaces—0
Distilling Spikes: Knowledge Distillation in Spiking Neural Networks—0
Distributed Low Precision Training Without Mixed Precision—0
Domain Generalization on Efficient Acoustic Scene Classification using Residual Normalization—0
Bias in Pruned Vision Models: In-Depth Analysis and Countermeasures—0
An Automatic and Efficient BERT Pruning for Edge AI Systems—0
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models—0
Analysis of Quantization on MLP-based Vision Models—0
AdaDeep: A Usage-Driven, Automated Deep Model Compression Framework for Enabling Ubiquitous Intelligent Mobiles—0
Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression Experiments—0
Beware of Calibration Data for Pruning Large Language Models—0
Analysis of memory consumption by neural networks based on hyperparameters—0
Benchmarking Adversarial Robustness of Compressed Deep Learning Models—0
An Algorithm-Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers—0
ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation—0
BD-KD: Balancing the Divergences for Online Knowledge Distillation—0
An Efficient Real-Time Object Detection Framework on Resource-Constricted Hardware Devices via Software and Hardware Co-design—0
Activation Sparsity Opportunities for Compressing General Large Language Models—0
Bayesian Federated Model Compression for Communication and Computation Efficiency—0
Bayesian Deep Learning Via Expectation Maximization and Turbo Deep Approximate Message Passing—0
A Model Compression Method with Matrix Product Operators for Speech Enhancement—0
A Mixed Integer Programming Approach for Verifying Properties of Binarized Neural Networks—0
Balancing Specialization, Generalization, and Compression for Detection and Tracking—0
Balancing Cost and Benefit with Tied-Multi Transformers—0
Activation Map Adaptation for Effective Knowledge Distillation—0
Single-path Bit Sharing for Automatic Loss-aware Model Compression—0
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications—0
Distilling Inductive Bias: Knowledge Distillation Beyond Model Compression—0
Extending DeepSDF for automatic 3D shape retrieval and similarity transform estimation—0
A Memory-Efficient Learning Framework for SymbolLevel Precoding with Quantized NN Weights—0
AMD: Automatic Multi-step Distillation of Large-scale Vision Models—0
Deep Model Compression Via Two-Stage Deep Reinforcement Learning—0
Deep Model Compression: Distilling Knowledge from Noisy Teachers—0
Deep Model Compression based on the Training History—0
A Web-Based Solution for Federated Learning with LLM-Based Automation—0
Discrete Model Compression With Resource Constraint for Deep Neural Networks—0
AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression Rates—0
Deep learning model compression using network sensitivity and gradients—0
Neural Epitome Search for Architecture-Agnostic Network Compression—0
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent—0
DeepRebirth: Accelerating Deep Neural Network Execution on Mobile Devices—0
DiPaCo: Distributed Path Composition—0
AMD: Adaptive Masked Distillation for Object Detection—0
Show:102550
← PrevPage 7 of 28Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MobileBERT + 2bit-1dim model compression using DKMAccuracy82.13—Unverified
2MobileBERT + 1bit-1dim model compression using DKMAccuracy63.17—Unverified