SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 551600 of 982 papers

TitleStatusHype
LeVoice ASR Systems for the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge0
Binaural Speech Enhancement Using STOI-Optimal Masks0
Speech Enhancement Using Self-Supervised Pre-Trained Model and Vector Quantization0
Speech Enhancement with Perceptually-motivated Optimization and Dual Transformations0
GIST-AiTeR System for the Diarization Task of the 2022 VoxCeleb Speaker Recognition Challenge0
A Universally-Deployable ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement, and Voice Separation0
Multimodal Speech Enhancement Using Burst Propagation0
Multi-View Attention Transfer for Efficient Speech Enhancement0
Speech Enhancement and Dereverberation with Diffusion-based Generative Models0
DNN-Free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online FastMNMF0
ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding0
Multi-channel target speech enhancement based on ERB-scaled spatial coherence features0
Improving spatial cues for hearables using a parameterized binaural CDR estimator0
Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational Environments0
Improving Visual Speech Enhancement Network by Learning Audio-visual Affinity with Multi-head Attention0
GLD-Net: Improving Monaural Speech Enhancement by Learning Global and Local Dependency Features with GLD Block0
Challenges and Opportunities in Multi-device Speech Processing0
SAQAM: Spatial Audio Quality Assessment Metric0
Efficient Transformer-based Speech Enhancement Using Long Frames and STFT Magnitudes0
Multi-channel end-to-end neural network for speech enhancement, source localization, and voice activity detection0
0/1 Deep Neural Networks via Block Coordinate Descent0
NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional Resampling0
EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal Learning0
Adversarial Privacy Protection on Speech EnhancementCode0
To Dereverb Or Not to Dereverb? Perceptual Studies On Real-Time Dereverberation Targets0
Canonical Cortical Graph Neural Networks and its Application for Speech Enhancement in Audio-Visual Hearing Aids0
Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus0
Joint Training of Speech Enhancement and Self-supervised Model for Noise-robust ASR0
NeuralEcho: A Self-Attentive Recurrent Neural Network For Unified Acoustic Echo Suppression And Speech Enhancement0
Dictionary-Based Fusion of Contact and Acoustic Microphones for Wind Noise Reduction0
Streaming Noise Context Aware Enhancement For Automatic Speech Recognition in Multi-Talker Environments0
Task splitting for DNN-based acoustic echo and noise removal0
A deep representation learning speech enhancement method using β-VAE0
Generalized Fast Multichannel Nonnegative Matrix Factorization Based on Gaussian Scale Mixtures for Blind Source Separation0
Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition0
Acoustic echo suppression using a learning-based multi-frame minimum variance distortionless response filter0
On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant training0
Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention0
A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network0
Improved far-field speech recognition using Joint Variational Autoencoder0
RadioSES: mmWave-Based Audioradio Speech Enhancement and Separation System0
Receptive Field Analysis of Temporal Convolutional Networks for Monaural Speech DereverberationCode0
Listen only to me! How well can target speech extraction handle false alarms?0
Exploiting Hidden Representations from a DNN-based Speech Recogniser for Speech Intelligibility Prediction in Hearing-impaired ListenersCode0
FFC-SE: Fast Fourier Convolution for Speech Enhancement0
Expression-preserving face frontalization improves visually assisted speech processing0
Complex Recurrent Variational Autoencoder with Application to Speech EnhancementCode0
Audio-visual multi-channel speech separation, dereverberation and recognition0
Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation0
End-to-End Integration of Speech Recognition, Speech Enhancement, and Self-Supervised Learning Representation0
Show:102550
← PrevPage 12 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99Unverified
2PESQetarianPESQ (wb)3.82Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7Unverified
5SEMamba (+PCS)PESQ (wb)3.69Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63Unverified
7PrimeK-NetPESQ (wb)3.61Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61Unverified
9MP-SENetPESQ (wb)3.6Unverified
10PCS_CS_WAVLMPESQ (wb)3.54Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4Unverified
2DTLNSI-SDR-WB16.34Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22Unverified
4ZipEnhancer (M)PESQ-WB3.81Unverified
5TF-Locoformer (M)PESQ-WB3.72Unverified
6ZipEnhancer (S)PESQ-WB3.69Unverified
7MambAttentionPESQ-WB3.67Unverified
8MP-SENetPESQ-WB3.62Unverified
9xLSTM-SENetPESQ-WB3.59Unverified
10BSRNN-S + MRSDPESQ-WB3.53Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67Unverified
2CA Dense U-Net (Complex)SDR18.64Unverified
3Dense U-Net (Complex)SDR18.4Unverified
4Dense U-Net (Real)SDR16.86Unverified
5U-Net (Real)SDR15.97Unverified
6Noisy/unprocessedSDR6.5Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09Unverified
2SGMSE+PESQ-WB2.5Unverified
3Demucs v4PESQ-WB2.37Unverified
4Schrödinger BridgePESQ-WB2.33Unverified
5Conv-TasNetPESQ-WB2.31Unverified
6CDiffuSEPESQ-WB1.6Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19Unverified
2ReVISE (bf)Audio Quality MOS4.11Unverified
3Demucs (ch2)Audio Quality MOS2.95Unverified
4Demucs (bf)Audio Quality MOS2.39Unverified
5MaxDI (Baseline)PESQ1.17Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08Unverified
2DCCRN-MCPESQ-NB3.21Unverified
3DCCRN-MPESQ-NB3.15Unverified
4DCCRNPESQ-NB3.04Unverified
5RNN-ModulationPESQ-WB2.75Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8Unverified
2SEMambaESTOI0.8Unverified
3xLSTM-SENetESTOI0.8Unverified
4MP-SENetESTOI0.79Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84Unverified
2DTLNPESQ2.23Unverified
3UnprocessedPESQ1.83Unverified
4Non-Real-Time MultiScale+PESQ1.52Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44Unverified
2DCCRN-MPESQ-NB3.28Unverified
3DCUNetPESQ-NB3.25Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82Unverified
2SpatialNetDNSMOS BAK3.43Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99Unverified
2ROSE-CDPESQ3.49Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07Unverified