SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 150 of 982 papers

TitleStatusHype
Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingCode9
Hybrid Transformers for Music Source SeparationCode5
TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorchCode4
DeepFilterNet2: Towards Real-Time Speech Enhancement on Embedded Devices for Full-Band AudioCode4
Deep Multi-Frame Filtering for Hearing AidsCode4
DeepFilterNet: Perceptually Motivated Real-Time Speech EnhancementCode4
Real-Time Packet Loss Concealment With Mixed Generative and Predictive ModelCode3
SoundStream: An End-to-End Neural Audio CodecCode3
An Investigation of Incorporating Mamba for Speech EnhancementCode3
SonicSim: A customizable simulation platform for speech processing in moving sound source scenariosCode3
Apollo: Band-sequence Modeling for High-Quality Audio RestorationCode3
Separate Anything You DescribeCode3
VoiceFixer: A Unified Framework for High-Fidelity Speech RestorationCode3
Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech SeparationCode3
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and DereverberationCode3
DeepFilterNet: A Low Complexity Speech Enhancement Framework for Full-Band Audio based on Deep FilteringCode2
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and EnhancementCode2
Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNetCode2
Speech Denoising in the Waveform Domain with Self-AttentionCode2
CMGAN: Conformer-Based Metric-GAN for Monaural Speech EnhancementCode2
StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and DereverberationCode2
Training-Free Multi-Step Audio Source SeparationCode2
Conditional Diffusion Probabilistic Model for Speech EnhancementCode2
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASRCode2
CMGAN: Conformer-based Metric GAN for Speech EnhancementCode2
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative SynchronizationCode2
Real Time Speech Enhancement in the Waveform DomainCode2
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase SpectraCode2
SEGAN: Speech Enhancement Generative Adversarial NetworkCode2
Mamba-SEUNet: Mamba UNet for Monaural Speech EnhancementCode2
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech EnhancementCode2
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech EnhancementCode2
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean SpeechCode2
FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech EnhancementCode2
ICASSP 2022 Acoustic Echo Cancellation ChallengeCode2
FlowSE: Efficient and High-Quality Speech Enhancement via Flow MatchingCode2
FSPEN: AN ULTRA-LIGHTWEIGHT NETWORK FOR REAL TIME SPEECH ENAHNCMENTCode2
ICASSP 2023 Acoustic Echo Cancellation ChallengeCode2
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPTCode2
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech EnhancementCode2
A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR ConditionsCode2
Mamba in Speech: Towards an Alternative to Self-AttentionCode2
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech EnhancementCode2
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency LossesCode2
Proximal Policy Optimization AlgorithmsCode2
Fast FullSubNet: Accelerate Full-band and Sub-band Fusion Model for Single-channel Speech EnhancementCode2
Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational EnvironmentsCode2
Efficient Speech Enhancement via Embeddings from Pre-trained Generative AudioencodersCode2
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow MatchingCode2
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTSCode2
Show:102550
← PrevPage 1 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified