SOTAVerified

Speech Enhancement

Speech Enhancement is a signal processing task that involves improving the quality of speech signals captured under noisy or degraded conditions. The goal of speech enhancement is to make speech signals clearer, more intelligible, and more pleasant to listen to, which can be used for various applications such as voice recognition, teleconferencing, and hearing aids. A representative Github project with online demo : ClearerVoice-Studio.

( Image credit: A Fully Convolutional Neural Network For Speech Enhancement )

Papers

Showing 301–350 of 982 papers

TitleStatusHype
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness—0
語音增強基於小腦模型控制器(A Speech Enhancement System Based on Cerebellar Model Articulation Controller) [In Chinese]—0
Combining Spatial Clustering with LSTM Speech Models for Multichannel Speech Enhancement—0
Collaborative Deep Learning for Speech Enhancement: A Run-Time Model Selection Method Using Autoencoders—0
A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction—0
Artificial Intelligence for Cochlear Implants: Review of Strategies, Challenges, and Perspectives—0
Fast Real-time Personalized Speech Enhancement: End-to-End Enhancement Network (E3Net) and Knowledge Distillation—0
FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network—0
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend—0
Cold Diffusion for Speech Enhancement—0
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model—0
Coarse-to-fine Optimization for Speech Enhancement—0
A scalable noisy speech dataset and online subjective test framework—0
Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions—0
Artifact-free Sound Quality in DNN-based Closed-loop Systems for Audio Processing—0
CleanUNet 2: A Hybrid Speech Denoising Model on Waveform and Spectrogram—0
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers—0
A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network—0
CLCNet: Deep learning-based Noise Reduction for Hearing Aids using Complex Linear Coding—0
Array Configuration-Agnostic Personal Voice Activity Detection Based on Spatial Coherence—0
CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings—0
Array Configuration-Agnostic Personalized Speech Enhancement using Long-Short-Term Spatial Coherence—0
Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement—0
CheapNET: Improving Light-weight speech enhancement network by projected loss function—0
Characterizing Speech Adversarial Examples Using Self-Attention U-Net Enhancement—0
A Robust Maximum Likelihood Distortionless Response Beamformer based on a Complex Generalized Gaussian Distribution—0
Challenges and Opportunities in Multi-device Speech Processing—0
ARiSE: Auto-Regressive Multi-Channel Speech Enhancement—0
A Low-Power Streaming Speech Enhancement Accelerator For Edge Devices—0
A deep representation learning speech enhancement method using β-VAE—0
A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation—0
Exploiting Time-Frequency Conformers for Music Audio Enhancement—0
Cellular Network Speech Enhancement: Removing Background and Transmission Noise—0
Face Recognition with Machine Learning in OpenCV_ Fusion of the results with the Localization Data of an Acoustic Camera for Speaker Identification—0
A Recurrent Variational Autoencoder for Speech Enhancement—0
Causal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech Enhancement—0
All Information is Necessary: Integrating Speech Positive and Negative Information by Contrastive Learning for Speech Enhancement—0
FADI-AEC: Fast Score Based Diffusion Model Guided by Far-end Signal for Acoustic Echo Cancellation—0
Evaluating the Intelligibility Benefits of Neural Speech Enrichment for Listeners with Normal Hearing and Hearing Impairment using the Greek Harvard Corpus—0
Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation—0
Can We Trust Deep Speech Prior?—0
Evaluating Speech Enhancement Systems Through Listening Effort—0
Can we steal your vocal identity from the Internet?: Initial investigation of cloning Obama's voice using GAN, WaveNet and low-quality found data—0
A Preliminary Study of the Application of Discrete Wavelet Transform Features in Conv-TasNet Speech Enhancement Model—0
ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding—0
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features—0
ESPnet-se: end-to-end speech enhancement and separation toolkit designed for asr integration—0
Exploiting the compressed spectral loss for the learning of the DEMUCS speech enhancement network—0
Canonical Cortical Graph Neural Networks and its Application for Speech Enhancement in Audio-Visual Hearing Aids—0
EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal Learning—0
Show:102550
← PrevPage 7 of 20Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ROSE-CD(PESQ)PESQ (wb)3.99—Unverified
2PESQetarianPESQ (wb)3.82—Unverified
3Mamba-SEUNet L (+PCS)PESQ (wb)3.73—Unverified
4Schrödinger bridge (PESQ loss)PESQ (wb)3.7—Unverified
5SEMamba (+PCS)PESQ (wb)3.69—Unverified
6ZipEnhancer (S, \lamba_6 = 0)PESQ (wb)3.63—Unverified
7PrimeK-NetPESQ (wb)3.61—Unverified
8ZipEnhancer (S, \lamba_6 = 0.2)PESQ (wb)3.61—Unverified
9MP-SENetPESQ (wb)3.6—Unverified
10PCS_CS_WAVLMPESQ (wb)3.54—Unverified
#ModelMetricClaimedVerifiedStatus
1BSRNN-S + MGDSI-SDR-WB21.4—Unverified
2DTLNSI-SDR-WB16.34—Unverified
3Non-Real-Time MultiScale+SI-SDR-WB16.22—Unverified
4ZipEnhancer (M)PESQ-WB3.81—Unverified
5TF-Locoformer (M)PESQ-WB3.72—Unverified
6ZipEnhancer (S)PESQ-WB3.69—Unverified
7MambAttentionPESQ-WB3.67—Unverified
8MP-SENetPESQ-WB3.62—Unverified
9xLSTM-SENetPESQ-WB3.59—Unverified
10BSRNN-S + MRSDPESQ-WB3.53—Unverified
#ModelMetricClaimedVerifiedStatus
1Inter-Channel Conv-TasNetSDR19.67—Unverified
2CA Dense U-Net (Complex)SDR18.64—Unverified
3Dense U-Net (Complex)SDR18.4—Unverified
4Dense U-Net (Real)SDR16.86—Unverified
5U-Net (Real)SDR15.97—Unverified
6Noisy/unprocessedSDR6.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Schrödinger Bridge (PESQ loss)PESQ-WB3.09—Unverified
2SGMSE+PESQ-WB2.5—Unverified
3Demucs v4PESQ-WB2.37—Unverified
4Schrödinger BridgePESQ-WB2.33—Unverified
5Conv-TasNetPESQ-WB2.31—Unverified
6CDiffuSEPESQ-WB1.6—Unverified
#ModelMetricClaimedVerifiedStatus
1ReVISE (ch2)Audio Quality MOS4.19—Unverified
2ReVISE (bf)Audio Quality MOS4.11—Unverified
3Demucs (ch2)Audio Quality MOS2.95—Unverified
4Demucs (bf)Audio Quality MOS2.39—Unverified
5MaxDI (Baseline)PESQ1.17—Unverified
6DAJA (MVDR,HMA,1000) (Overlapped Speech)SDR-4.76—Unverified
#ModelMetricClaimedVerifiedStatus
1ZipEnhancer (M)PESQ-NB4.08—Unverified
2DCCRN-MCPESQ-NB3.21—Unverified
3DCCRN-MPESQ-NB3.15—Unverified
4DCCRNPESQ-NB3.04—Unverified
5RNN-ModulationPESQ-WB2.75—Unverified
#ModelMetricClaimedVerifiedStatus
1MambAttentionESTOI0.8—Unverified
2SEMambaESTOI0.8—Unverified
3xLSTM-SENetESTOI0.8—Unverified
4MP-SENetESTOI0.79—Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ2.84—Unverified
2DTLNPESQ2.23—Unverified
3UnprocessedPESQ1.83—Unverified
4Non-Real-Time MultiScale+PESQ1.52—Unverified
#ModelMetricClaimedVerifiedStatus
1DCUNet-MCPESQ-NB3.44—Unverified
2DCCRN-MPESQ-NB3.28—Unverified
3DCUNetPESQ-NB3.25—Unverified
#ModelMetricClaimedVerifiedStatus
1CleanMel-L-mapDNSMOS3.82—Unverified
2SpatialNetDNSMOS BAK3.43—Unverified
#ModelMetricClaimedVerifiedStatus
1rose_cd(PESQ )PESQ3.99—Unverified
2ROSE-CDPESQ3.49—Unverified
#ModelMetricClaimedVerifiedStatus
1Wave-U-NetCBAK3.24—Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ2.7—Unverified
#ModelMetricClaimedVerifiedStatus
1SE-MelGANAudio Quality MOS3.1—Unverified
#ModelMetricClaimedVerifiedStatus
1DeFT-ANPESQ3.01—Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refPESQ3.03—Unverified
#ModelMetricClaimedVerifiedStatus
1SepFormerPESQ3.07—Unverified