SOTAVerified

Model Compression

Model Compression is an actively pursued area of research over the last few years with the goal of deploying state-of-the-art deep networks in low-power and resource limited devices without significant drop in accuracy. Parameter pruning, low-rank factorization and weight quantization are some of the proposed methods to compress the size of deep networks.

Source: KD-MRI: A knowledge distillation framework for image reconstruction and image restoration in MRI workflow

Papers

Showing 451500 of 1356 papers

TitleStatusHype
EoRA: Training-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation0
CompMarkGS: Robust Watermarking for Compressed 3D Gaussian Splatting0
Ensemble-Compression: A New Method for Parallel Training of Deep Neural Networks0
Enhancing Targeted Attack Transferability via Diversified Weight Pruning0
Complexity-Driven CNN Compression for Resource-constrained Edge AI0
Architecture Compression0
Compacting Deep Neural Networks for Internet of Things: Methods and Applications0
Enhanced Sparsification via Stimulative Training0
CompactifAI: Extreme Compression of Large Language Models using Quantum-Inspired Tensor Networks0
Energy-efficient Knowledge Distillation for Spiking Neural Networks0
EncCluster: Scalable Functional Encryption in Federated Learning through Weight Clustering and Probabilistic Filters0
Compact CNN Structure Learning by Knowledge Distillation0
A Progressive Sub-Network Searching Framework for Dynamic Inference0
A Deep Cascade Network for Unaligned Face Attribute Classification0
Accelerating Machine Learning Primitives on Commodity Hardware0
Enabling Deep Learning on Edge Devices through Filter Pruning and Knowledge Transfer0
Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications0
Communication-Efficient Federated Learning with Adaptive Compression under Dynamic Bandwidth0
Empowering Edge Intelligence: A Comprehensive Survey on On-Device AI Models0
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks0
Communication-Efficient Distributed Online Learning with Kernels0
A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration Framework0
E-LANG: Energy-Based Joint Inferencing of Super and Swift Language Models0
Efficient Transformer Knowledge Distillation: A Performance Review0
Approximability and Generalisation0
Enabling All In-Edge Deep Learning: A Literature Review0
Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression0
Efficient Speech Representation Learning with Low-Bit Quantization0
Efficient Recurrent Neural Networks using Structured Matrices in FPGAs0
CoLLD: Contrastive Layer-to-layer Distillation for Compressing Multilingual Pre-trained Speech Encoders0
Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy0
Energy-Efficient Model Compression and Splitting for Collaborative Inference Over Time-Varying Channels0
Additive Tree-Structured Covariance Function for Conditional Parameter Spaces in Bayesian Optimization0
Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured Sparsity0
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion0
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations0
Efficient Point Cloud Classification via Offline Distillation Framework and Negative-Weight Self-Distillation Technique0
Collaborative Teacher-Student Learning via Multiple Knowledge Transfer0
Efficient Neural Networks for Tiny Machine Learning: A Comprehensive Review0
Efficient Model Compression Techniques with FishLeg0
Towards Feature Distribution Alignment and Diversity Enhancement for Data-Free Quantization0
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome0
Applications of Knowledge Distillation in Remote Sensing: A Survey0
Error-aware Quantization through Noise Tempering0
ADC/DAC-Free Analog Acceleration of Deep Neural Networks with Frequency Transformation0
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models0
Efficient Model Compression for Hierarchical Federated Learning0
Efficient Model Compression for Bayesian Neural Networks0
Efficient Memory Management for GPU-based Deep Learning Systems0
ClusComp: A Simple Paradigm for Model Compression and Efficient Finetuning0
Show:102550
← PrevPage 10 of 28Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MobileBERT + 2bit-1dim model compression using DKMAccuracy82.13Unverified
2MobileBERT + 1bit-1dim model compression using DKMAccuracy63.17Unverified