SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 101150 of 982 papers

TitleStatusHype
BASPRO: a balanced script producer for speech corpus collection based on the genetic algorithmCode1
SpeechLMScore: Evaluating speech generation using speech language modelCode1
High Fidelity Speech Enhancement with Band-split RNNCode1
McNet: Fuse Multiple Cues for Multichannel Speech EnhancementCode1
SceneFake: An Initial Dataset and Benchmarks for Scene Fake Audio DetectionCode1
Inference and Denoise: Causal Inference-based Neural Speech EnhancementCode1
Diffusion-based Generative Speech Source SeparationCode1
Diffiner: A Versatile Diffusion-based Generative Refiner for Speech EnhancementCode1
MMS-MSG: A Multi-purpose Multi-Speaker Mixture Signal GeneratorCode1
Improving Speech Enhancement through Fine-Grained Speech CharacteristicsCode1
A light-weight full-band speech enhancement modelCode1
Insights Into Deep Non-linear Filters for Improved Multi-channel Speech EnhancementCode1
A Systematic Comparison of Phonetic Aware Techniques for Speech EnhancementCode1
On the Role of Spatial, Spectral, and Temporal Processing for DNN-based Non-linear Multi-channel Speech EnhancementCode1
Universal Speech Enhancement with Score-based DiffusionCode1
U-Former: Improving Monaural Speech Enhancement with Multi-head Self and Cross AttentionCode1
Boosting Self-Supervised Embeddings for Speech EnhancementCode1
Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-SynthesisCode1
Perceptual Contrast Stretching on Target Feature for Speech EnhancementCode1
Speech Enhancement with Score-Based Generative Models in the Complex STFT DomainCode1
Dual-Path Style Learning for End-to-End Noise-Robust Speech RecognitionCode1
HiFi++: a Unified Framework for Bandwidth Extension and Speech EnhancementCode1
MANNER: Multi-view Attention Network for Noise ErasureCode1
Look\&Listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech EnhancementCode1
L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office EnvironmentCode1
RemixIT: Continual self-training of speech enhancement models via bootstrapped remixingCode1
HGCN: Harmonic gated compensation network for speech enhancementCode1
Towards Intelligibility-Oriented Audio-Visual Speech EnhancementCode1
MultiSV: Dataset for Far-Field Multi-Channel Speaker VerificationCode1
Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal TransportCode1
Uformer: A Unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberationCode1
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain FeaturesCode1
Continual self-training with bootstrapped remixing for speech enhancementCode1
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingCode1
Toward Degradation-Robust Voice ConversionCode1
Dual-branch Attention-In-Attention Transformer for single-channel speech enhancementCode1
MetricGAN-U: Unsupervised speech enhancement/ dereverberation based only on noisy/ reverberated speechCode1
Interactive Feature Fusion for End-to-End Noise-Robust Speech RecognitionCode1
NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesCode1
A Deep Learning Loss Function based on Auditory Power Compression for Speech EnhancementCode1
Complex-valued Spatial Autoencoders for Multichannel Speech EnhancementCode1
A Causal U-net based Neural Beamforming Network for Real-Time Multi-Channel Speech EnhancementCode1
Microphone Array Generalization for Multichannel Narrowband Deep Speech EnhancementCode1
A Study on Speech Enhancement Based on Diffusion Probabilistic ModelCode1
Multi-Task Audio Source SeparationCode1
EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy EnvironmentsCode1
TENET: A Time-reversal Enhancement Network for Noise-robust ASRCode1
Unsupervised Speech Enhancement using Dynamical Variational Auto-EncodersCode1
MeshRIR: A Dataset of Room Impulse Responses on Meshed Grid Points For Evaluating Sound Field Analysis and Synthesis MethodsCode1
Attention-based distributed speech enhancement for unconstrained microphone arrays with varying number of nodesCode1
Show:102550
← PrevPage 3 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified