SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 301350 of 982 papers

TitleStatusHype
Target Speech Extraction with Conditional Diffusion Model0
Efficient Monaural Speech Enhancement using Spectrum Attention Fusion0
SAMbA: Speech enhancement with Asynchronous ad-hoc Microphone Arrays0
PCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement0
The Effect of Spoken Language on Speech Enhancement using Self-Supervised Speech Representation Loss FunctionsCode0
Single Channel Speech Enhancement Using U-Net Spiking Neural NetworksCode0
Non Intrusive Intelligibility Predictor for Hearing Impaired Individuals using Self Supervised Speech Representations0
MetricGAN-OKD: Multi-Metric Optimization of MetricGAN via Online Knowledge Distillation for Speech EnhancementCode1
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs0
Low bit rate binaural link for improved ultra low-latency low-complexity multichannel speech enhancement in Hearing Aids0
Noise-aware Speech Enhancement using Diffusion Probabilistic ModelCode1
Audio-Visual Speech Enhancement Using Self-supervised Learning to Improve Speech Intelligibility in Cochlear Implant Simulations0
Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition0
Disentanglement in a GAN for Unconditional Speech SynthesisCode1
Multi-Loss Convolutional Network with Time-Frequency Attention for Speech Enhancement0
Feature Normalization for Fine-tuning Self-Supervised Models in Speech Enhancement0
Variance-Preserving-Based Interpolation Diffusion Models for Speech EnhancementCode1
Unsupervised speech enhancement with deep dynamical generative speech and noise models0
Audio-Visual Speech Enhancement With Selective Off-Screen Speech Extraction0
Efficient Encoder-Decoder and Dual-Path Conformer for Comprehensive Feature Learning in Speech Enhancement0
Convolutional Recurrent Neural Network with Attention for 3D Speech Enhancement0
A Mask Free Neural Network for Monaural Speech EnhancementCode1
On the Behavior of Intrusive and Non-intrusive Speech Enhancement Metrics in Predictive and Generative Settings0
EffCRN: An Efficient Convolutional Recurrent Network for High-Performance Speech Enhancement0
Influence of Lossy Speech Codecs on Hearing-aid, Binaural Sound Source Localisation using DNNs0
On Crowdsourcing-design with Comparison Category Rating for Evaluating Speech Enhancement Algorithms0
Audio-Visual Speech Enhancement with Score-Based Generative Models0
Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model0
A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint ModelsCode1
Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation0
Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss0
SE-Bridge: Speech Enhancement with Consistent Brownian Bridge0
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase SpectraCode2
DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting0
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders0
Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found DataCode1
BASEN: Time-Domain Brain-Assisted Speech Enhancement Network with Convolutional Cross Attention in Multi-talker ConditionsCode1
Integrating Uncertainty into Neural Network-based Speech EnhancementCode1
Deep Multi-Frame Filtering for Hearing AidsCode4
DeepFilterNet: Perceptually Motivated Real-Time Speech EnhancementCode4
Diffusion-based Signal Refiner for Speech Separation0
All Information is Necessary: Integrating Speech Positive and Negative Information by Contrastive Learning for Speech Enhancement0
Neural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated Full- and Sub-Band Modeling0
Array Configuration-Agnostic Personal Voice Activity Detection Based on Spatial Coherence0
The future of hearing aid technology0
Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance FieldsCode2
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR0
Attention-based Speech Enhancement Using Human Quality Perception Modelling0
A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI0
Transformers in Speech Processing: A Survey0
Show:102550
← PrevPage 7 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified