SOTAVerified

Model Compression

Model Compression is an actively pursued area of research over the last few years with the goal of deploying state-of-the-art deep networks in low-power and resource limited devices without significant drop in accuracy. Parameter pruning, low-rank factorization and weight quantization are some of the proposed methods to compress the size of deep networks.

Source: KD-MRI: A knowledge distillation framework for image reconstruction and image restoration in MRI workflow

Papers

Showing 251–300 of 1356 papers

TitleStatusHype
Comprehensive Survey of Model Compression and Speed up for Vision Transformers—0
Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices—0
Compressed models are NOT miniature versions of large models—0
Artemis: HE-Aware Training for Efficient Privacy-Preserving Machine Learning—0
Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models—0
A Novel Architecture Slimming Method for Network Pruning and Knowledge Distillation—0
Adaptive Learning of Tensor Network Structures—0
Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models—0
Accelerating Framework of Transformer by Hardware Design and Model Compression Co-Optimization—0
Deep Model Compression based on the Training History—0
Channel Compression: Rethinking Information Redundancy among Channels in CNN Architecture—0
DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization—0
An Improving Framework of regularization for Network Compression—0
Order of Compression: A Systematic and Optimal Sequence to Combinationally Compress CNN—0
Adaptive Quantization of Neural Networks—0
Deep learning model compression using network sensitivity and gradients—0
Deep Model Compression: Distilling Knowledge from Noisy Teachers—0
Accelerating deep neural networks for efficient scene understanding in automotive cyber-physical systems—0
Adaptive Neural Connections for Sparsity Learning—0
Deep Collective Knowledge Distillation—0
Cascaded channel pruning using hierarchical self-distillation—0
Can We Find Strong Lottery Tickets in Generative Models?—0
A New Clustering-Based Technique for the Acceleration of Deep Convolutional Networks—0
Can Students Outperform Teachers in Knowledge Distillation based Model Compression?—0
Can Students Beyond The Teacher? Distilling Knowledge from Teacher's Bias—0
A "Network Pruning Network" Approach to Deep Model Compression—0
An Empirical Study of Low Precision Quantization for TinyML—0
Can Model Compression Improve NLP Fairness—0
Heterogeneous Federated Learning using Dynamic Model Pruning and Adaptive Gradient—0
2-bit Model Compression of Deep Convolutional Neural Network on ASIC Engine for Image Retrieval—0
Deep Compression of Neural Networks for Fault Detection on Tennessee Eastman Chemical Processes—0
Deep Model Compression Via Two-Stage Deep Reinforcement Learning—0
Can collaborative learning be private, robust and scalable?—0
CAIT: Triple-Win Compression towards High Accuracy, Fast Inference, and Favorable Transferability For ViTs—0
Multihop: Leveraging Complex Models to Learn Accurate Simple Models—0
Bringing AI To Edge: From Deep Learning's Perspective—0
An Empirical Investigation of Matrix Factorization Methods for Pre-trained Transformers—0
Adapting Models to Signal Degradation using Distillation—0
BRIEDGE: EEG-Adaptive Edge AI for Multi-Brain to Multi-Robot Interaction—0
Bridging the Resource Gap: Deploying Advanced Imitation Learning Models onto Affordable Embedded Platforms—0
A Multi-objective Complex Network Pruning Framework Based on Divide-and-conquer and Global Performance Impairment Ranking—0
Bridging the Gap Between Foundation Models and Heterogeneous Federated Learning—0
An Embedded Deep Learning Object Detection Model For Traffic In Asian Countries—0
AdapMTL: Adaptive Pruning Framework for Multitask Learning Model—0
Accelerating Deep Learning with Dynamic Data Pruning—0
Debiased Distillation by Transplanting the Last Layer—0
Boosting Graph Neural Networks via Adaptive Knowledge Distillation—0
Block-wise Intermediate Representation Training for Model Compression—0
An Efficient Sparse Inference Software Accelerator for Transformer-based Language Models on CPUs—0
Block Skim Transformer for Efficient Question Answering—0
Show:102550
← PrevPage 6 of 28Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MobileBERT + 2bit-1dim model compression using DKMAccuracy82.13—Unverified
2MobileBERT + 1bit-1dim model compression using DKMAccuracy63.17—Unverified