SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 51100 of 982 papers

TitleStatusHype
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech EnhancementCode2
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech EnhancementCode2
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean SpeechCode2
Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNetCode2
VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram MaskingCode2
Integrating Uncertainty into Neural Network-based Speech EnhancementCode1
Improving Speech Enhancement through Fine-Grained Speech CharacteristicsCode1
Inference and Denoise: Causal Inference-based Neural Speech EnhancementCode1
Interactive Feature Fusion for End-to-End Noise-Robust Speech RecognitionCode1
AnCoGen: Analysis, Control and Generation of Speech with a Masked AutoencoderCode1
Insights Into Deep Non-linear Filters for Improved Multi-channel Speech EnhancementCode1
Instantaneous PSD Estimation for Speech Enhancement based on Generalized Principal ComponentsCode1
Improved Lite Audio-Visual Speech EnhancementCode1
Improving GANs for Speech EnhancementCode1
Improving Perceptual Quality by Phone-Fortified Perceptual Loss using Wasserstein Distance for Speech EnhancementCode1
INTERSPEECH 2021 ConferencingSpeech Challenge: Towards Far-field Multi-Channel Speech Enhancement for Video ConferencingCode1
HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial NetworksCode1
HiFi++: a Unified Framework for Bandwidth Extension and Speech EnhancementCode1
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial NetworksCode1
Group Communication with Context Codec for Lightweight Source SeparationCode1
Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech RecognitionCode1
HGCN: Harmonic gated compensation network for speech enhancementCode1
High Fidelity Speech Enhancement with Band-split RNNCode1
A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint ModelsCode1
FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial NetworksCode1
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech EnhancementCode1
A Modulation-Domain Loss for Neural-Network-based Real-time Speech EnhancementCode1
A Study on Speech Enhancement Based on Diffusion Probabilistic ModelCode1
A Causal U-net based Neural Beamforming Network for Real-Time Multi-Channel Speech EnhancementCode1
Hold Me Tight: Stable Encoder-Decoder Design for Speech EnhancementCode1
Investigating the Design Space of Diffusion Models for Speech EnhancementCode1
Explainable DNN-based Beamformer with PostfilterCode1
A Mask Free Neural Network for Monaural Speech EnhancementCode1
EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy EnvironmentsCode1
Disentanglement in a GAN for Unconditional Speech SynthesisCode1
Dual-branch Attention-In-Attention Transformer for single-channel speech enhancementCode1
Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found DataCode1
Diffusion-based Generative Speech Source SeparationCode1
Dual-Path Style Learning for End-to-End Noise-Robust Speech RecognitionCode1
DNN-based mask estimation for distributed speech enhancement in spatially unconstrained microphone arraysCode1
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative ConditionsCode1
A Refining Underlying Information Framework for Monaural Speech EnhancementCode1
A Differentiable Perceptual Audio Metric Learned from Just Noticeable DifferencesCode1
Deep Residual-Dense Lattice Network for Speech EnhancementCode1
DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech EnhancementCode1
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup ProcessingCode1
A Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband SpeechCode1
Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance MatricesCode1
An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and SeparationCode1
A light-weight full-band speech enhancement modelCode1
Show:102550
← PrevPage 2 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified