SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 951–982 of 982 papers

TitleStatusHype
EMGSE: Acoustic/EMG Fusion for Multimodal Speech Enhancement—0
Employing low-pass filtered temporal speech features for the training of ideal ratio mask in speech enhancement—0
End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression—0
End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation—0
End-to-End Model for Speech Enhancement by Consistent Spectrogram Masking—0
End-to-End Neural Speech Coding for Real-Time Communications—0
End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural Networks—0
Enhancement and Recognition of Reverberant and Noisy Speech by Extending Its Coherence—0
Enhancement of Noisy Speech Exploiting an Exponential Model Based Threshold and a Custom Thresholding Function in Perceptual Wavelet Packet Domain—0
Enhancement of Noisy Speech with Low Speech Distortion Based on Probabilistic Geometric Spectral Subtraction—0
Enhancement of Spatial Clustering-Based Time-Frequency Masks using LSTM Neural Networks—0
Enhancing Speech Quality through the Integration of BGRU and Transformer Architectures—0
EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal Learning—0
ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration—0
ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding—0
Evaluating Speech Enhancement Systems Through Listening Effort—0
Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation—0
Evaluating the Intelligibility Benefits of Neural Speech Enrichment for Listeners with Normal Hearing and Hearing Impairment using the Greek Harvard Corpus—0
Exploiting the compressed spectral loss for the learning of the DEMUCS speech enhancement network—0
Exploiting Time-Frequency Conformers for Music Audio Enhancement—0
Exploration of Adapter for Noise Robust Automatic Speech Recognition—0
Exploring Length Generalization For Transformer-based Speech Enhancement—0
Multi-Channel Speaker Verification for Single and Multi-talker Speech—0
Exploring Speech Enhancement for Low-resource Speech Synthesis—0
Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition—0
Exploring the Best Loss Function for DNN-Based Low-latency Speech Enhancement with Temporal Convolutional Networks—0
Exploring the Potential of Data-Driven Spatial Audio Enhancement Using a Single-Channel Model—0
Exploring WavLM on Speech Enhancement—0
Expression-preserving face frontalization improves visually assisted speech processing—0
Face Recognition with Machine Learning in OpenCV_ Fusion of the results with the Localization Data of an Acoustic Camera for Speaker Identification—0
FADI-AEC: Fast Score Based Diffusion Model Guided by Far-end Signal for Acoustic Echo Cancellation—0
Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus—0
Show:102550
← PrevPage 20 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99—Unverified
2PESQetarianPESQ (wb)3.82—Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73—Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7—Unverified
5SEMamba (+PCS)PESQ (wb)3.69—Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63—Unverified
7PrimeK-NetPESQ (wb)3.61—Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61—Unverified
9MP-SENetPESQ (wb)3.6—Unverified
10PCS_CS_WAVLMPESQ (wb)3.54—Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4—Unverified
2DTLNSI-SDR-WB16.34—Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22—Unverified
4ZipEnhancer (M)PESQ-WB3.81—Unverified
5TF-Locoformer (M)PESQ-WB3.72—Unverified
6ZipEnhancer (S)PESQ-WB3.69—Unverified
7MambAttentionPESQ-WB3.67—Unverified
8MP-SENetPESQ-WB3.62—Unverified
9xLSTM-SENetPESQ-WB3.59—Unverified
10BSRNN-S + MRSDPESQ-WB3.53—Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67—Unverified
2CA Dense U-Net (Complex)SDR18.64—Unverified
3Dense U-Net (Complex)SDR18.4—Unverified
4Dense U-Net (Real)SDR16.86—Unverified
5U-Net (Real)SDR15.97—Unverified
6Noisy/unprocessedSDR6.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09—Unverified
2SGMSE+PESQ-WB2.5—Unverified
3Demucs v4PESQ-WB2.37—Unverified
4Schrödinger BridgePESQ-WB2.33—Unverified
5Conv-TasNetPESQ-WB2.31—Unverified
6CDiffuSEPESQ-WB1.6—Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19—Unverified
2ReVISE (bf)Audio Quality MOS4.11—Unverified
3Demucs (ch2)Audio Quality MOS2.95—Unverified
4Demucs (bf)Audio Quality MOS2.39—Unverified
5MaxDI (Baseline)PESQ1.17—Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76—Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08—Unverified
2DCCRN-MCPESQ-NB3.21—Unverified
3DCCRN-MPESQ-NB3.15—Unverified
4DCCRNPESQ-NB3.04—Unverified
5RNN-ModulationPESQ-WB2.75—Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8—Unverified
2SEMambaESTOI0.8—Unverified
3xLSTM-SENetESTOI0.8—Unverified
4MP-SENetESTOI0.79—Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84—Unverified
2DTLNPESQ2.23—Unverified
3UnprocessedPESQ1.83—Unverified
4Non-Real-Time MultiScale+PESQ1.52—Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44—Unverified
2DCCRN-MPESQ-NB3.28—Unverified
3DCUNetPESQ-NB3.25—Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82—Unverified
2SpatialNetDNSMOS BAK3.43—Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99—Unverified
2ROSE-CDPESQ3.49—Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24—Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7—Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1—Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01—Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03—Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07—Unverified