SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 401450 of 982 papers

TitleStatusHype
Generative Pre-training for Speech with Flow Matching0
LC-TTFS: Towards Lossless Network Conversion for Spiking Neural Networks with TTFS Coding0
Deep Beamforming for Speech Enhancement and Speaker Localization with an Array Response-Aware Loss Function0
Real-time Speech Enhancement and Separation with a Unified Deep Neural Network for Single/Dual Talker Scenarios0
A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction0
Magnitude-and-phase-aware Speech Enhancement with Parallel Sequence Modeling0
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms0
VSANet: Real-time Speech Enhancement Based on Voice Activity Detection and Causal Spatial Attention0
An experiment on an automated literature survey of data-driven speech enhancement methods0
An Exploration of Task-decoupling on Two-stage Neural Post Filter for Real-time Personalized Acoustic Echo Cancellation0
MBTFNet: Multi-Band Temporal-Frequency Neural Network For Singing Voice Enhancement0
uSee: Unified Speech Enhancement and Editing with Conditional Diffusion Models0
A Fused Deep Denoising Sound Coding Strategy for Bilateral Cochlear Implants0
Toward Universal Speech Enhancement for Diverse Input Conditions0
Multichannel Voice Trigger Detection Based on Transform-average-concatenate0
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study0
DDTSE: Discriminative Diffusion Model for Target Speech Extraction0
AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data0
Speech enhancement with frequency domain auto-regressive modeling0
A Multiscale Autoencoder (MSAE) Framework for End-to-End Neural Network Speech Enhancement0
Deep Complex U-Net with Conformer for Audio-Visual Speech Enhancement0
Joint Minimum Processing Beamforming and Near-end Listening Enhancement0
Posterior sampling algorithms for unsupervised speech enhancement with recurrent variational autoencoder0
Exploring Speech Enhancement for Low-resource Speech Synthesis0
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement0
Diffusion-based speech enhancement with a weighted generative-supervised learning loss0
Refining DNN-based Mask Estimation using CGMM-based EM Algorithm for Multi-channel Noise Reduction0
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning0
Unifying Robustness and Fidelity: A Comprehensive Study of Pretrained Generative Methods for Speech Enhancement in Adverse Conditions0
Two-Step Knowledge Distillation for Tiny Speech Enhancement0
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement0
Assessing the Generalization Gap of Learning-Based Speech Enhancement Systems in Noisy and Reverberant Environments0
CleanUNet 2: A Hybrid Speech Denoising Model on Waveform and Spectrogram0
PlumberNet: Fixing interference leakage after GEV beamformingCode0
Spiking Structured State Space Model for Monaural Speech Enhancement0
Causal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech Enhancement0
Single-Channel Speech Enhancement with Deep Complex U-Networks and Probabilistic Latent Space Models0
Noise robust speech emotion recognition with signal-to-noise ratio adapting speech enhancement0
Rep2wav: Noise Robust text-to-speech Using self-supervised representations0
Exploiting Time-Frequency Conformers for Music Audio Enhancement0
AdVerb: Visually Guided Audio Dereverberation0
Convoifilter: A case study of doing cocktail party speech recognition0
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer0
Target Speech Extraction with Conditional Diffusion Model0
Efficient Monaural Speech Enhancement using Spectrum Attention Fusion0
SAMbA: Speech enhancement with Asynchronous ad-hoc Microphone Arrays0
PCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement0
The Effect of Spoken Language on Speech Enhancement using Self-Supervised Speech Representation Loss FunctionsCode0
Single Channel Speech Enhancement Using U-Net Spiking Neural NetworksCode0
Non Intrusive Intelligibility Predictor for Hearing Impaired Individuals using Self Supervised Speech Representations0
Show:102550
← PrevPage 9 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified