SOTAVerified

Speech Separation

The task of extracting all overlapping speech sources in a given mixed speech signal refers to the Speech Separation. Speech Separation is a special scenario of source separation problem, where the focus is only on the overlapping speech signal sources and other interferences such as music or noise signals are not the main concern of the study. A recent representative Github project can be referred to ClearerVoice-Studio.

Source: A Unified Framework for Speech Separation

Image credit: Speech Separation of A Target Speaker Based on Deep Neural Networks

Papers

Showing 301350 of 359 papers

TitleStatusHype
Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement0
Onssen: an open-source speech separation and enhancement libraryCode0
Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation0
Interrupted and cascaded permutation invariant training for speech separationCode0
Mixup-breakdown: a consistency training method for improving generalization of speech separation models0
A Multi-Phase Gammatone Filterbank for Speech Separation via TasNetCode0
Analyzing the impact of speaker localization errors on speech separation for automatic speech recognitionCode0
Multi-channel Speech Separation Using Deep Embedding Model with Multilayer Bootstrap Networks0
Filterbank design for end-to-end speech separationCode0
Two-Step Sound Source Separation: Training on Learned Latent TargetsCode0
Multi-Talker MVDR Beamforming Based on Extended Complex Gaussian Mixture Model0
MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition0
Probabilistic Permutation Invariant Training for Speech Separation0
Discriminative Learning for Monaural Speech Separation Using Deep Embedding Features0
WHAM!: Extending Speech Separation to Noisy EnvironmentsCode0
Single-Channel Speech Separation with Auxiliary Speaker Embeddings0
A comprehensive study of speech separation: spectrogram vs waveform separation0
End-to-End Multi-Channel Speech Separation0
Universal Sound Separation0
Divide and Conquer: A Deep CASA Approach to Talker-independent Monaural Speaker SeparationCode0
Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and ClusteringCode0
Low-Latency Speaker-Independent Continuous Speech Separation0
Orthonormal Embedding-based Deep Clustering for Single-channel Speech Separation0
Tensor-Train Long Short-Term Memory for Monaural Speech Enhancement0
Semi-Supervised Monaural Singing Voice Separation With a Masking Network Trained on Synthetic MixturesCode0
Face Landmark-based Speaker-Independent Audio-Visual Speech Enhancement in Multi-Talker EnvironmentsCode0
Building Corpora for Single-Channel Speech Separation Across Multiple Domains0
Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures using Spatial InformationCode0
End-to-End Monaural Multi-speaker ASR System without Pretraining0
Recognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural Networks0
End-to-end Networks for Supervised Single-channel Speech Separation0
Real-time Single-channel Dereverberation and Separation with Time-domainAudio Separation NetworkCode0
DNN driven Speaker Independent Audio-Visual Mask Estimation for Speech Separation0
Sound Signal Processing with Seq2Tree Network0
End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction0
Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech SeparationCode0
Alternative Objective Functions for Deep ClusteringCode0
The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines0
TasNet: time-domain audio separation network for real-time, single-channel speech separationCode0
Singing Voice Separation with Deep U-Net Convolutional NetworksCode0
Deep Recurrent NMF for Speech Separation by Unfolding Iterative ThresholdingCode0
Using Optimal Ratio Mask as Training Target for Supervised Speech Separation0
Supervised Speech Separation Based on Deep Learning: An Overview0
Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition0
Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training0
Speaker-independent Speech Separation with Deep Attractor Network0
Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural NetworksCode0
Deep attractor network for single-microphone speaker separationCode0
Deep Clustering and Conventional Networks for Music Separation: Stronger Together0
Monaural Multi-Talker Speech Recognition using Factorial Speech Processing Models0
Show:102550
← PrevPage 7 of 8Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1SepReformer-LSI-SDRi25.1Unverified
2TF-Locoformer (L) + DMSI-SDRi25.1Unverified
3TF-Locoformer (M) + DMSI-SDRi24.6Unverified
4TF-Locoformer (L)SI-SDRi24.2Unverified
5MossFormer2 (L)SI-SDRi24.1Unverified
6SepTDA (L=12)SI-SDRi24Unverified
7Separate And DiffuseSI-SDRi23.9Unverified
8TF-Locoformer (M)SI-SDRi23.6Unverified
9TF-Locoformer (S) + DMSI-SDRi22.8Unverified
10MossFormer (L) + DMSI-SDRi22.8Unverified
#ModelMetricClaimedVerifiedStatus
1TF-Locoformer (M)SI-SDRi18.5Unverified
2TF-Locoformer (S)SI-SDRi17.4Unverified
3SepReformer-L + DMSI-SDRi17.1Unverified
4MossFormer2SI-SDRi17Unverified
5MossFormer (L) + DMSI-SDRi16.3Unverified
6TD-Conformer (XL) + DMSI-SDRi14.6Unverified
7Improved Sudo rm -rf (U=36)SI-SDRi13.5Unverified
8TD-Conformer (L) + DMSI-SDRi13.4Unverified
9WavesplitSI-SDRi13.2Unverified
10DPTNET - SRSSNSI-SDRi12.3Unverified
#ModelMetricClaimedVerifiedStatus
1MossFormer2 (w speed perturb)SI-SDRi22.2Unverified
2TF-Locoformer (M)SI-SDRi22.1Unverified
3MossFormer2 (w/o DM)SI-SDRi21.7Unverified
4Separate And DiffuseSI-SDRi21.5Unverified
5WHYVSI-SDRi17.5Unverified
6TDANet LargeSI-SDRi17.4Unverified
7TDANetSI-SDRi16.9Unverified
8Conv-Tasnet (Libri1Mix speech enhancement pre-trained)SI-SDRi14.1Unverified
9Conv-Tasnet (Libri1Mix speech enhancement multi-task)SI-SDRi13.7Unverified
10Conv-TasnetSI-SDRi13.2Unverified
#ModelMetricClaimedVerifiedStatus
1SepTDASI-SDRi23.7Unverified
2MossFormer2SI-SDRi22.2Unverified
3MossFormer (L) + DMSI-SDRi21.2Unverified
4Separate And DiffuseSI-SDRi20.9Unverified
5MossFormer (M) + DMSI-SDRi20.8Unverified
6SepItSI-SDRi20.1Unverified
7SepFormerSI-SDRi19.5Unverified
8SandglassetSI-SDRi17.1Unverified
9Gated DualPathRNNSI-SDRi16.85Unverified
#ModelMetricClaimedVerifiedStatus
1IIANetSI-SNRi16.4Unverified
2TDFNet-largeSI-SNRi15.8Unverified
3TDFNet (MHSA + Shared)SI-SNRi15Unverified
4RTFS-Net-12SI-SNRi14.9Unverified
5RTFS-Net-6SI-SNRi14.6Unverified
6CTCNetSI-SNRi14.3Unverified
7RTFS-Net-4SI-SNRi14.1Unverified
8TDFNet-smallSI-SNRi13.6Unverified
#ModelMetricClaimedVerifiedStatus
1SepReformer-L + DMSI-SDRi18.4Unverified
2MossFormer2SI-SDRi18.1Unverified
3MossFormer (L) + DMSI-SDRi17.3Unverified
4TDANet LargeSI-SDRi15.2Unverified
5TDANetSI-SDRi14.8Unverified
6WHYVSI-SDRi12.96Unverified
#ModelMetricClaimedVerifiedStatus
1SepTDASI-SDRi21Unverified
2Hungarian PITSI-SDRi13.22Unverified
3Conditional TasNetSI-SDRi11.7Unverified
4TasTasSI-SDRi11.14Unverified
5Gated DualPathRNNSI-SDRi10.56Unverified
6Multi-Decoder DPRNNSI-SDRi5.9Unverified
#ModelMetricClaimedVerifiedStatus
1IIANetSI-SNRi18.3Unverified
2RTFS-Net-12SI-SNRi17.5Unverified
3CTCNetSI-SNRi17.4Unverified
4RTFS-Net-6SI-SNRi16.9Unverified
5RTFS-Net-4SI-SNRi15.5Unverified
#ModelMetricClaimedVerifiedStatus
1IIANetSI-SNRi14Unverified
2RTFS-Net-12SI-SNRi12.4Unverified
3CTCNetSI-SNRi11.9Unverified
4RTFS-Net-6SI-SNRi11.8Unverified
5RTFS-Net-4SI-SNRi11.5Unverified
#ModelMetricClaimedVerifiedStatus
1SepTDASI-SDRi22Unverified
2Gated DualPathRNNSI-SDRi12.88Unverified
3Conditional TasNetSI-SDRi12.5Unverified
4OR-PITSI-SDRi10.2Unverified
5Multi-Decoder DPRNNSI-SDRi9.3Unverified
#ModelMetricClaimedVerifiedStatus
1Separate And DiffuseSI-SDRi14.2Unverified
2SepItSI-SDRi13.7Unverified
3OCDSI-SDRi13.4Unverified
4Hungarian PITSI-SDRi12.72Unverified
#ModelMetricClaimedVerifiedStatus
1Separate And DiffuseSI-SDRi9Unverified
2SepItSI-SDRi8.2Unverified
3Hungarian PITSI-SDRi7.78Unverified
#ModelMetricClaimedVerifiedStatus
1SDR9.6Unverified
2Audio-Visual concat-refSDR8.05Unverified
#ModelMetricClaimedVerifiedStatus
1Separate And DiffuseSI-SDRi5.2Unverified
2Hungarian PITSI-SDRi4.26Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer (base)0S5.6Unverified
2Conformer (large)0S5.4Unverified
#ModelMetricClaimedVerifiedStatus
1Hungarian PITSI-SDRi5.66Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refSDR10.55Unverified
#ModelMetricClaimedVerifiedStatus
1MossFormer2SI-SDRi20.5Unverified