SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 401450 of 982 papers

TitleStatusHype
McNet: Fuse Multiple Cues for Multichannel Speech EnhancementCode1
Array Configuration-Agnostic Personalized Speech Enhancement using Long-Short-Term Spatial Coherence0
Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-channel Speech Enhancement0
Hybrid Transformers for Music Source SeparationCode5
Multi-Label Training for Text-Independent Speaker Identification0
The Potential of Neural Speech Synthesis-based Data Augmentation for Personalized Speech Enhancement0
SceneFake: An Initial Dataset and Benchmarks for Scene Fake Audio DetectionCode1
Cross-Attention is all you need: Real-Time Streaming Transformers for Personalised Speech Enhancement0
DiffPhase: Generative Diffusion-based STFT Phase Retrieval0
Egocentric Audio-Visual Noise Suppression0
Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation0
Self-Supervised Learning for Speech Enhancement through SynthesisCode0
Analysing Diffusion-based Generative Approaches versus Discriminative Approaches for Speech Restoration0
Speech enhancement using ego-noise references with a microphone array embedded in an unmanned aerial vehicle0
Real-Time Joint Personalized Speech Enhancement and Acoustic Echo Cancellation0
Cold Diffusion for Speech Enhancement0
Iterative autoregression: a novel trick to improve your low-latency speech enhancement model0
Dynamic Kernels and Channel Attention for Low Resource Speaker Verification0
Fast and efficient speech enhancement with variational autoencoders0
A weighted-variance variational autoencoder model for speech enhancement0
Analysis of Noisy-target Training for DNN-based speech enhancement0
Inference and Denoise: Causal Inference-based Neural Speech EnhancementCode1
Audio-visual speech enhancement with a deep Kalman filter generative model0
Exploiting the compressed spectral loss for the learning of the DEMUCS speech enhancement network0
A Preliminary Study of the Application of Discrete Wavelet Transform Features in Conv-TasNet Speech Enhancement Model0
SCA: Streaming Cross-attention Alignment for Echo Cancellation0
Audio-Visual Speech Enhancement and Separation by Utilizing Multi-Modal Self-Supervised Embeddings0
Diffusion-based Generative Speech Source SeparationCode1
Diffiner: A Versatile Diffusion-based Generative Refiner for Speech EnhancementCode1
A Training and Inference Strategy Using Noisy and Enhanced Speech as Target for Speech Enhancement without Clean SpeechCode0
Parallel Gated Neural Network With Attention Mechanism For Speech Enhancement0
SCP-GAN: Self-Correcting Discriminator Optimization for Training Consistency Preserving Metric GAN on Speech Enhancement Tasks0
TridentSE: Guiding Speech Enhancement with 32 Global Tokens0
Time-Domain Speech Enhancement for Robust Automatic Speech Recognition0
A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids0
Improved Normalizing Flow-Based Speech Enhancement using an All-pole Gammatone Filterbank for Conditional Input Representation0
spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid filtering for multi-channel speech enhancement0
Accelerating RNN-based Speech Enhancement on a Multi-Core MCU with Mixed FP16-INT8 Post-Training Quantization0
LeVoice ASR Systems for the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge0
Binaural Speech Enhancement Using STOI-Optimal Masks0
Speech Enhancement Using Self-Supervised Pre-Trained Model and Vector Quantization0
Speech Enhancement with Perceptually-motivated Optimization and Dual Transformations0
MMS-MSG: A Multi-purpose Multi-Speaker Mixture Signal GeneratorCode1
CMGAN: Conformer-Based Metric-GAN for Monaural Speech EnhancementCode2
GIST-AiTeR System for the Diarization Task of the 2022 VoxCeleb Speaker Recognition Challenge0
A Universally-Deployable ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement, and Voice Separation0
Multimodal Speech Enhancement Using Burst Propagation0
Multi-View Attention Transfer for Efficient Speech Enhancement0
Speech Enhancement and Dereverberation with Diffusion-based Generative Models0
DNN-Free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online FastMNMF0
Show:102550
← PrevPage 9 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified