SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 801850 of 982 papers

TitleStatusHype
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement0
A two-step backward compatible fullband speech enhancement system0
A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI0
Audio Recording Device Identification Based on Deep Learning0
Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition0
Audio-visual multi-channel speech separation, dereverberation and recognition0
Audio-Visual Speech Enhancement and Separation by Utilizing Multi-Modal Self-Supervised Embeddings0
Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks0
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues0
Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks0
Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders0
Audio-Visual Speech Enhancement Using Self-supervised Learning to Improve Speech Intelligibility in Cochlear Implant Simulations0
Audio-Visual Speech Enhancement With Selective Off-Screen Speech Extraction0
Audio-visual speech enhancement with a deep Kalman filter generative model0
Audio-Visual Speech Enhancement with Score-Based Generative Models0
A Unified Deep Learning Framework for Short-Duration Speaker Verification in Adverse Environments0
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model0
A Universally-Deployable ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement, and Voice Separation0
AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data0
Autoregressive Speech Enhancement via Acoustic Tokens0
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement0
AV Speech Enhancement Challenge using a Real Noisy Corpus0
Batch-normalized joint training for DNN-based distant speech recognition0
BERT for Joint Multichannel Speech Dereverberation with Spatial-aware Tasks0
Binaural Speech Enhancement Using STOI-Optimal Masks0
Blind Acoustic Room Parameter Estimation Using Phase Features0
Blind Mask to Improve Intelligibility of Non-Stationary Noisy Speech0
Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training0
Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing0
Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation0
Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic Modeling0
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement0
Building a Luganda Text-to-Speech Model From Crowdsourced Data0
Building state-of-the-art distant speech recognition using the CHiME-4 challenge with a setup of speech enhancement baseline0
Canonical Cortical Graph Neural Networks and its Application for Speech Enhancement in Audio-Visual Hearing Aids0
Can we steal your vocal identity from the Internet?: Initial investigation of cloning Obama's voice using GAN, WaveNet and low-quality found data0
Can We Trust Deep Speech Prior?0
Causal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech Enhancement0
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features0
Cellular Network Speech Enhancement: Removing Background and Transmission Noise0
Challenges and Opportunities in Multi-device Speech Processing0
Characterizing Speech Adversarial Examples Using Self-Attention U-Net Enhancement0
CheapNET: Improving Light-weight speech enhancement network by projected loss function0
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings0
CLCNet: Deep learning-based Noise Reduction for Hearing Aids using Complex Linear Coding0
CleanUNet 2: A Hybrid Speech Denoising Model on Waveform and Spectrogram0
Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions0
Coarse-to-fine Optimization for Speech Enhancement0
Cold Diffusion for Speech Enhancement0
Collaborative Deep Learning for Speech Enhancement: A Run-Time Model Selection Method Using Autoencoders0
Show:102550
← PrevPage 17 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified