SOTAVerified

Model Compression

Model Compression is an actively pursued area of research over the last few years with the goal of deploying state-of-the-art deep networks in low-power and resource limited devices without significant drop in accuracy. Parameter pruning, low-rank factorization and weight quantization are some of the proposed methods to compress the size of deep networks.

Source: KD-MRI: A knowledge distillation framework for image reconstruction and image restoration in MRI workflow

Papers

Showing 851–900 of 1356 papers

TitleStatusHype
Toward Extremely Low Bit and Lossless Accuracy in DNNs with Progressive ADMM—0
Model Compression via Hyper-Structure Network—0
Model Compression via Symmetries of the Parameter Space—0
Toward Real-World Voice Disorder Classification—0
Model Compression with Generative Adversarial Networks—0
Model Compression with Multi-Task Knowledge Distillation for Web-scale Question Answering System—0
Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System—0
An Effective Information Theoretic Framework for Channel Pruning—0
Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification—0
On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks—0
Towards Accurate Post-Training Quantization for Vision Transformer—0
A Light-weight Deep Human Activity Recognition Algorithm Using Multi-knowledge Distillation—0
Towards a tailored mixed-precision sub-8-bit quantization scheme for Gated Recurrent Units using Genetic Algorithms—0
Modular Transformers: Compressing Transformers into Modularized Layers for Flexible Efficient Inference—0
Modulating Regularization Frequency for Efficient Compression-Aware Model Training—0
MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness—0
MPruner: Optimizing Neural Network Size with CKA-Based Mutual Information Pruning—0
MSP: An FPGA-Specific Mixed-Scheme, Multi-Precision Deep Neural Network Quantization Framework—0
MT-BioNER: Multi-task Learning for Biomedical Named Entity Recognition using Deep Bidirectional Transformers—0
Towards Better Parameter-Efficient Fine-Tuning for Large Language Models: A Position Paper—0
Multi-Dimensional Pruning: A Unified Framework for Model Compression—0
Towards Building a Real Time Mobile Device Bird Counting System Through Synthetic Data Training and Model Compression—0
Multi-head Knowledge Distillation for Model Compression—0
An Automatic and Efficient BERT Pruning for Edge AI Systems—0
Towards domain generalisation in ASR with elitist sampling and ensemble knowledge distillation—0
Multi-Precision Quantized Neural Networks via Encoding Decomposition of -1 and +1—0
MultiPruner: Balanced Structure Removal in Foundation Models—0
Multi-stage Progressive Compression of Conformer Transducer for On-device Speech Recognition—0
Multi-task Learning Approach for Modulation and Wireless Signal Classification for 5G and Beyond: Edge Deployment via Model Compression—0
Multi-Task Semantic Communications via Large Models—0
Multi-Task Zipping via Layer-wise Neuron Sharing—0
MWQ: Multiscale Wavelet Quantized Neural Networks—0
N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning—0
Analysis of Quantization on MLP-based Vision Models—0
N-Ary Quantization for CNN Model Compression and Inference Acceleration—0
NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search—0
Natively Interpretable Machine Learning and Artificial Intelligence: Preliminary Results and Future Directions—0
NeR-VCP: A Video Content Protection Method Based on Implicit Neural Representation—0
Reconstructing Pruned Filters using Cheap Spatial Transformations—0
Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining—0
Network Pruning for Low-Rank Binary Index—0
Network Pruning for Low-Rank Binary Indexing—0
Weight Normalization based Quantization for Deep Neural Network Compression—0
Neural 3D Scene Compression via Model Compression—0
Neural Architecture Codesign for Fast Bragg Peak Analysis—0
ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation—0
Neural Network Compression for Noisy Storage Devices—0
Neural Network Compression using Binarization and Few Full-Precision Weights—0
Neural Network Compression Via Sparse Optimization—0
Neural Network Pruning by Cooperative Coevolution—0
Show:102550
← PrevPage 18 of 28Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MobileBERT + 2bit-1dim model compression using DKMAccuracy82.13—Unverified
2MobileBERT + 1bit-1dim model compression using DKMAccuracy63.17—Unverified