SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 51100 of 982 papers

TitleStatusHype
Linguistic Knowledge Transfer Learning for Speech Enhancement0
ProSE: Diffusion Priors for Speech Enhancement0
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture SearchCode2
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech EnhancementCode2
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASRCode2
PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel Convolutional Neural Networks for Single Channel Speech EnhancementCode1
Enhancing Speech Quality through the Integration of BGRU and Transformer Architectures0
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec0
Adaptive Convolution for CNN-based Speech Enhancement ModelsCode1
LMFCA-Net: A Lightweight Model for Multi-Channel Speech Enhancement with Efficient Narrow-Band and Cross-Band Attention0
TAPS: Throat and Acoustic Paired Speech Dataset for Deep Learning-Based Speech Enhancement0
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge0
Advances in Microphone Array Processing and Multichannel Speech Enhancement0
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling0
Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingCode9
Learning-based A Posteriori Speech Presence Probability Estimation and Applications0
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement0
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement0
UP-Cycle-SENet: Unpaired Phase-aware Speech Enhancement Using Deep Complex Cycle Adversarial Networks0
Let SSMs be ConvNets: State-space Modeling with Optimal Tensor ContractionsCode0
Speech Enhancement with Overlapped-Frame Information Fusion and Causal Self-AttentionCode0
SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts AggregationCode1
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids0
Microphone Array Signal Processing and Deep Learning for Speech Enhancement0
Multi-modal Speech Enhancement with Limited Electromyography Channels0
Estimation and Restoration of Unknown Nonlinear Distortion using DiffusionCode0
xLSTM-SENet: xLSTM for Single-Channel Speech EnhancementCode2
AnCoGen: Analysis, Control and Generation of Speech with a Masked AutoencoderCode1
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow MatchingCode2
Artifact-free Sound Quality in DNN-based Closed-loop Systems for Audio Processing0
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features0
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario0
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology0
Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech EnhancementCode0
Scalable Speech Enhancement with Dynamic Channel Pruning0
Mamba-SEUNet: Mamba UNet for Monaural Speech EnhancementCode2
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling0
Investigating the Effects of Diffusion-based Conditional Generative Speech Models Used for Speech Enhancement on Dysarthric Speech0
Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation0
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch0
Source Separation & Automatic Transcription for MusicCode1
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation0
Towards Advanced Speech Signal Processing: A Statistical Perspective on Convolution-Based Architectures and its Applications0
GhostRNN: Reducing State Redundancy in RNN with Cheap Operations0
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions0
Explainable DNN-based Beamformer with PostfilterCode1
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model0
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions0
Selective State Space Model for Monaural Speech Enhancement0
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement0
Show:102550
← PrevPage 2 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified