SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 851900 of 982 papers

TitleStatusHype
Combining Spatial Clustering with LSTM Speech Models for Multichannel Speech Enhancement0
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness0
Comparative Study between Adversarial Networks and Classical Techniques for Speech Enhancement0
Comparison of remote experiments using crowdsourcing and laboratory experiments on speech intelligibility0
Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network for Speech Enhancement0
Complex spectrogram enhancement by convolutional neural network with multi-metrics learning0
Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker Verification0
Consistency-aware multi-channel speech enhancement using deep neural networks0
Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy0
Contextual Audio-Visual Switching For Speech Enhancement in Real-World Environments0
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning0
Controlling the Perceived Sound Quality for Dialogue Enhancement with Deep Learning0
Convoifilter: A case study of doing cocktail party speech recognition0
Convolutional Neural Network-based Speech Enhancement for Cochlear Implant Recipients0
Convolutional-Recurrent Neural Networks for Speech Enhancement0
ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion0
Cooperative Dual Attention for Audio-Visual Speech Enhancement with Facial Cues0
Correlating Subword Articulation with Lip Shapes for Embedding Aware Audio-Visual Speech Enhancement0
Cross-attention conformer for context modeling in speech enhancement for ASR0
Cross-Attention is all you need: Real-Time Streaming Transformers for Personalised Speech Enhancement0
Cross-domain Single-channel Speech Enhancement Model with Bi-projection Fusion Module for Noise-robust ASR0
ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement0
Cycle-Consistent Speech Enhancement0
D²Net: A Denoising and Dereverberation Network Based on Two-branch Encoder and Dual-path Transformer0
DASB -- Discrete Audio and Speech Benchmark0
Dataset of Spatial Room Impulse Responses in a Variable Acoustics Room for Six Degrees-of-Freedom Rendering and Analysis0
DCCRGAN: Deep Complex Convolution Recurrent Generator Adversarial Network for Speech Enhancement0
DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement0
DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting0
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions0
DDS: A new device-degraded speech dataset for speech enhancement0
Deep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation0
Deep Beamforming for Speech Enhancement and Speaker Localization with an Array Response-Aware Loss Function0
Deep Complex U-Net with Conformer for Audio-Visual Speech Enhancement0
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration0
Deep Interaction between Masking and Mapping Targets for Single-Channel Speech Enhancement0
Deep-Learning-Based Audio-Visual Speech Enhancement in Presence of Lombard Effect0
Deep Learning Based Speech Beamforming0
Deep learning for minimum mean-square error approaches to speech enhancement0
Deep low-latency joint speech transmission and enhancement over a gaussian channel0
Deep neural network Based Low-latency Speech Separation with Asymmetric analysis-Synthesis Window Pair0
Deep neural network techniques for monaural speech enhancement: state of the art analysis0
Deep Noise Suppression Maximizing Non-Differentiable PESQ Mediated by a Non-Intrusive PESQNet0
Deep Noise Suppression With Non-Intrusive PESQNet Supervision Enabling the Use of Real Training Data0
Deep Residual Echo Suppression and Noise Reduction: A Multi-Input FCRN Approach in a Hybrid Speech Enhancement System0
Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks0
Deep Time Delay Neural Network for Speech Enhancement with Full Data Learning0
Deep Unfolding: Model-Based Inspiration of Novel Deep Architectures0
Deep Xi as a Front-End for Robust Automatic Speech Recognition0
Dense CNN with Self-Attention for Time-Domain Speech Enhancement0
Show:102550
← PrevPage 18 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified