SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 551600 of 982 papers

TitleStatusHype
Improving Speech Recognition on Noisy Speech via Speech Enhancement with Multi-Discriminators CycleGAN0
Learning-based personal speech enhancement for teleconferencing by exploiting spatial-spectral features0
Harmonic and non-Harmonic Based Noisy Reverberant Speech Enhancement in Time Domain0
A Training Framework for Stereo-Aware Speech Enhancement using Deep Neural Networks0
使用低通時序列語音特徵訓練理想比率遮罩法之語音強化 (Employing Low-Pass Filtered Temporal Speech Features for the Training of Ideal Ratio Mask in Speech Enhancement)0
Effect of noise suppression losses on speech distortion and ASR performance0
Dataset of Spatial Room Impulse Responses in a Variable Acoustics Room for Six Degrees-of-Freedom Rendering and Analysis0
Towards Intelligibility-Oriented Audio-Visual Speech EnhancementCode1
A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation0
BLOOM-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech EnhancementCode0
Unsupervised Speech Enhancement with speech recognition embedding and disentanglement losses0
S-DCCRN: Super Wide Band DCCRN with learnable complex feature for speech enhancement0
Joint Far- and Near-End Speech Intelligibility Enhancement based on the Approximated Speech Intelligibility Index0
Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal TransportCode1
Uformer: A Unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberationCode1
MultiSV: Dataset for Far-Field Multi-Channel Speaker VerificationCode1
OSSEM: one-shot speaker adaptive speech enhancement using meta learning0
Inter-channel Conv-TasNet for multichannel speech enhancement0
SEOFP-NET: Compression and Acceleration of Deep Neural Networks for Speech Enhancement Using Sign-Exponent-Only Floating-Points0
Deep Noise Suppression Maximizing Non-Differentiable PESQ Mediated by a Non-Intrusive PESQNet0
Weight, Block or Unit? Exploring Sparsity Tradeoffs for Speech Enhancement on Tiny Neural Accelerators0
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain FeaturesCode1
Reduction of Subjective Listening Effort for TV Broadcast Signals with Recurrent Neural Networks0
SNRi Target Training for Joint Speech Enhancement and Recognition0
Cross-attention conformer for context modeling in speech enhancement for ASR0
Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction0
Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions0
One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement0
Continual self-training with bootstrapped remixing for speech enhancementCode1
Speech Enhancement-assisted Voice Conversion in Noisy Environments0
Speech Enhancement Based on Cyclegan with Noise-informed Training0
Personalized Speech Enhancement: New Models and Comprehensive Evaluation0
Similarity-and-Independence-Aware Beamformer with Iterative Casting and Boost Start for Target Source Extraction Using Reference0
Toward Degradation-Robust Voice ConversionCode1
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingCode1
Dual-branch Attention-In-Attention Transformer for single-channel speech enhancementCode1
Improving Character Error Rate Is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-box Acoustic Models0
MetricGAN-U: Unsupervised speech enhancement/ dereverberation based only on noisy/ reverberated speechCode1
DeepFilterNet: A Low Complexity Speech Enhancement Framework for Full-Band Audio based on Deep FilteringCode2
Wav2vec-Switch: Contrastive Learning from Original-noisy Speech Pairs for Robust Speech Recognition0
Interactive Feature Fusion for End-to-End Noise-Robust Speech RecognitionCode1
Aura: Privacy-preserving Augmentation to Improve Test Set Diversity in Speech EnhancementCode0
Lightweight Speech Enhancement in Unseen Noisy and Reverberant Conditions using KISS-GEV Beamforming0
PL-EESR: Perceptual Loss Based END-TO-END Robust Speaker Representation ExtractionCode0
End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression0
Employing low-pass filtered temporal speech features for the training of ideal ratio mask in speech enhancement0
Speech-MLP: a simple MLP architecture for speech processing0
Masks Fusion with Multi-Target Learning For Speech EnhancementCode0
NORESQA: A Framework for Speech Quality Assessment using Non-Matching ReferencesCode1
DDS: A new device-degraded speech dataset for speech enhancement0
Show:102550
← PrevPage 12 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified