SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 901950 of 982 papers

TitleStatusHype
Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement0
Design and Optimization of a Speech Recognition Front-End for Distant-Talking Control of a Music Playback Device0
DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing0
DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement0
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids0
DFSNet: A Steerable Neural Beamformer Invariant to Microphone Array Configuration for Real-Time, Low-Latency Speech Enhancement0
Dictionary-Based Fusion of Contact and Acoustic Microphones for Wind Noise Reduction0
Dictionary Update for NMF-based Voice Conversion Using an Encoder-Decoder Network0
DiffPhase: Generative Diffusion-based STFT Phase Retrieval0
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement0
Diffusion-based Signal Refiner for Speech Separation0
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler0
Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule0
Diffusion-based speech enhancement with a weighted generative-supervised learning loss0
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders0
Diffusion-based Unsupervised Audio-visual Speech Enhancement0
Diffusion Buffer: Online Diffusion-based Speech Enhancement with Sub-Second Latency0
DDTSE: Discriminative Diffusion Model for Target Speech Extraction0
Diffusion Models for Audio Restoration0
Dilated U-net based approach for multichannel speech enhancement from First-Order Ambisonics recordings0
Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational Environments0
Distributed Microphone Speech Enhancement based on Deep Learning0
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers0
DNN-Based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone Arrays0
DNN-Based Speech Presence Probability Estimation for Multi-Frame Single-Microphone Speech Enhancement0
DNN-Free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online FastMNMF0
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study0
Does Speech enhancement of publicly available data help build robust Speech Recognition Systems?0
Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss0
DPATD: Dual-Phase Audio Transformer for Denoising0
DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement0
An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions0
Dual-Stage Low-Complexity Reconfigurable Speech Enhancement0
Dynamic Acoustic Compensation and Adaptive Focal Training for Personalized Speech Enhancement0
Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement0
Dynamic Kernels and Channel Attention for Low Resource Speaker Verification0
EDNet: A Distortion-Agnostic Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training0
非負矩陣分解法於語音調變頻譜強化之研究(A study of enhancing the modulation spectrum of speech signals via nonnegative matrix factorization)[In Chinese]0
EffCRN: An Efficient Convolutional Recurrent Network for High-Performance Speech Enhancement0
Effect of noise suppression losses on speech distortion and ASR performance0
Effects of Lombard Reflex on the Performance of Deep-Learning-Based Audio-Visual Speech Enhancement Systems0
A Dual-Staged Context Aggregation Method Towards Efficient End-To-End Speech Enhancement0
Efficient Encoder-Decoder and Dual-Path Conformer for Comprehensive Feature Learning in Speech Enhancement0
Efficient High-Performance Bark-Scale Neural Network for Residual Echo and Noise Suppression0
Efficient Low-Latency Speech Enhancement with Mobile Audio Streaming Networks0
Efficient Monaural Speech Enhancement using Spectrum Attention Fusion0
Efficient Trainable Front-Ends for Neural Speech Enhancement0
Efficient Transformer-based Speech Enhancement Using Long Frames and STFT Magnitudes0
Egocentric Audio-Visual Noise Suppression0
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams0
Show:102550
← PrevPage 19 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified