SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 351400 of 982 papers

TitleStatusHype
A study on speech enhancement using exponent-only floating point quantized neural network (EOFP-QNN)0
Controlling the Perceived Sound Quality for Dialogue Enhancement with Deep Learning0
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning0
Multimodal Audio-Visual Information Fusion using Canonical-Correlated Graph Neural Network for Energy-Efficient Speech Enhancement0
Advances in Microphone Array Processing and Multichannel Speech Enhancement0
Acoustic echo suppression using a learning-based multi-frame minimum variance distortionless response filter0
Contextual Audio-Visual Switching For Speech Enhancement in Real-World Environments0
A Study of Incorporating Articulatory Movement Information in Speech Enhancement0
Constrained Convolutional-Recurrent Networks to Improve Speech Quality with Low Impact on Recognition Accuracy0
A Study of Enhancement, Augmentation, and Autoencoder Methods for Domain Adaptation in Distant Speech Recognition0
Consistency-aware multi-channel speech enhancement using deep neural networks0
Conditional Generative Adversarial Networks for Speech Enhancement and Noise-Robust Speaker Verification0
A Statistically Principled and Computationally Efficient Approach to Speech Enhancement using Variational Autoencoders0
Assessing the Generalization Gap of Learning-Based Speech Enhancement Systems in Noisy and Reverberant Environments0
A Monaural Speech Enhancement Method for Robust Small-Footprint Keyword Spotting0
Advanced Clustering Techniques for Speech Signal Enhancement: A Review and Metanalysis of Fuzzy C-Means, K-Means, and Kernel Fuzzy C-Means Methods0
Complex spectrogram enhancement by convolutional neural network with multi-metrics learning0
Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network for Speech Enhancement0
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations0
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling0
Generative Pre-training for Speech with Flow Matching0
Comparison of remote experiments using crowdsourcing and laboratory experiments on speech intelligibility0
A Speech Intelligibility Enhancement Model based on Canonical Correlation and Deep Learning for Hearing-Assistive Technologies0
Comparative Study between Adversarial Networks and Classical Techniques for Speech Enhancement0
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness0
語音增強基於小腦模型控制器(A Speech Enhancement System Based on Cerebellar Model Articulation Controller) [In Chinese]0
Combining Spatial Clustering with LSTM Speech Models for Multichannel Speech Enhancement0
Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement0
Collaborative Deep Learning for Speech Enhancement: A Run-Time Model Selection Method Using Autoencoders0
A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction0
A Model Compression Method with Matrix Product Operators for Speech Enhancement0
Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives0
A consolidated view of loss functions for supervised deep learning-based speech enhancement0
Accelerating RNN-based Speech Enhancement on a Multi-Core MCU with Mixed FP16-INT8 Post-Training Quantization0
Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement0
Frequency Gating: Improved Convolutional Neural Networks for Speech Enhancement in the Time-Frequency Domain0
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning0
Gated Recurrent Fusion with Joint Training Framework for Robust End-to-End Speech Recognition0
Generalized Fast Multichannel Nonnegative Matrix Factorization Based on Gaussian Scale Mixtures for Blind Source Separation0
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement0
Cold Diffusion for Speech Enhancement0
Generative Speech Enhancement Based on Cloned Networks0
Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition0
Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement0
GhostRNN: Reducing State Redundancy in RNN with Cheap Operations0
GIST-AiTeR System for the Diarization Task of the 2022 VoxCeleb Speaker Recognition Challenge0
GLD-Net: Improving Monaural Speech Enhancement by Learning Global and Local Dependency Features with GLD Block0
French Listening Tests for the Assessment of Intelligibility, Quality, and Identity of Body-Conducted Speech Enhancement0
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model0
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching0
Show:102550
← PrevPage 8 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified