SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 601650 of 982 papers

TitleStatusHype
Incorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems0
Time Alignment using Lip Images for Frame-based Electrolaryngeal Voice Conversion0
Machine Learning: Challenges, Limitations, and Compatibility for Audio Restoration Processes0
A Two-stage Complex Network using Cycle-consistent Generative Adversarial Networks for Speech Enhancement0
Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement0
Task-aware Warping Factors in Mask-based Speech Enhancement0
A Deep Learning Loss Function based on Auditory Power Compression for Speech EnhancementCode1
Cross-domain Single-channel Speech Enhancement Model with Bi-projection Fusion Module for Noise-robust ASR0
Deep Residual Echo Suppression and Noise Reduction: A Multi-Input FCRN Approach in a Hybrid Speech Enhancement System0
Complex-valued Spatial Autoencoders for Multichannel Speech EnhancementCode1
A Causal U-net based Neural Beamforming Network for Real-Time Multi-Channel Speech EnhancementCode1
Microphone Array Generalization for Multichannel Narrowband Deep Speech EnhancementCode1
Inplace Gated Convolutional Recurrent Neural Network For Dual-channel Speech Enhancement0
A Study on Speech Enhancement Based on Diffusion Probabilistic ModelCode1
Multitask-Based Joint Learning Approach To Robust ASR For Radio Communication Speech0
Controlling the Perceived Sound Quality for Dialogue Enhancement with Deep Learning0
Multi-Task Audio Source SeparationCode1
EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy EnvironmentsCode1
Incorporating Multi-Target in Multi-Stage Speech Enhancement Model for Better Generalization0
SoundStream: An End-to-End Neural Audio CodecCode3
TENET: A Time-reversal Enhancement Network for Noise-robust ASRCode1
DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement0
SRIB-LEAP submission to Far-field Multi-Channel Speech Enhancement Challenge for Video Conferencing0
Unsupervised Speech Enhancement using Dynamical Variational Auto-EncodersCode1
Deep neural network Based Low-latency Speech Separation with Asymmetric analysis-Synthesis Window Pair0
MeshRIR: A Dataset of Room Impulse Responses on Meshed Grid Points For Evaluating Sound Field Analysis and Synthesis MethodsCode1
A Flow-Based Neural Network for Time Domain Speech Enhancement0
DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement0
Attention-based distributed speech enhancement for unconstrained microphone arrays with varying number of nodesCode1
Learning Audio-Visual DereverberationCode1
Deep Interaction between Masking and Mapping Targets for Single-Channel Speech Enhancement0
Human Listening and Live Captioning: Multi-Task Training for Speech Enhancement0
Should We Always Separate?: Switching Between Enhanced and Observed Signals for Overlapping Speech Recognition0
A Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer And Large Scale Synthetic Data0
Phoneme-Based Ratio Mask Estimation for Reverberant Speech Enhancement in Cochlear Implant Processors0
An Improved Measure of Musical Noise Based on Spectral Kurtosis0
Training Speech Enhancement Systems with Noisy Speech Datasets0
RNNoise-Ex: Hybrid Speech Enhancement System based on RNN and Spectral FeaturesCode1
Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech EnhancementCode0
A time-domain nearfield frequency-invariant beamforming method0
Dual-Stage Low-Complexity Reconfigurable Speech Enhancement0
Separate but Together: Unsupervised Federated Learning for Speech Enhancement from Non-IID DataCode1
Test-Time Adaptation Toward Personalized Speech Enhancement: Zero-Shot Learning with Knowledge Distillation0
Zero-Shot Personalized Speech Enhancement through Speaker-Informed Model Selection0
Speech Enhancement using Separable Polling Attention and Global Layer Normalization followed with PReLU0
dEchorate: a Calibrated Room Impulse Response Database for Echo-aware Signal ProcessingCode1
Nonlinear Spatial Filtering in Multichannel Speech Enhancement0
Comparison of remote experiments using crowdsourcing and laboratory experiments on speech intelligibility0
L3DAS21 Challenge: Machine Learning for 3D Audio Signal ProcessingCode1
Complex Spectral Mapping With Attention Based Convolution Recurrent Neural Network for Speech Enhancement0
Show:102550
← PrevPage 13 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified