SOTAVerified

Speech Separation

The task of extracting all overlapping speech sources in a given mixed speech signal refers to the Speech Separation. Speech Separation is a special scenario of source separation problem, where the focus is only on the overlapping speech signal sources and other interferences such as music or noise signals are not the main concern of the study. A recent representative Github project can be referred to ClearerVoice-Studio.

Source: A Unified Framework for Speech Separation

Image credit: Speech Separation of A Target Speaker Based on Deep Neural Networks

Papers

Showing 251–300 of 359 papers

TitleStatusHype
Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-order Latent Domain—0
Streaming Target-Speaker ASR with Neural Transducer—0
Streaming Multi-talker Speech Recognition with Joint Speaker Identification—0
Study of the Performance of CEEMDAN in Underdetermined Speech Separation—0
Supervised Speech Separation Based on Deep Learning: An Overview—0
Surrogate Source Model Learning for Determined Source Separation—0
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer—0
TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024—0
Target Confusion in End-to-end Speaker Extraction: Analysis and Approaches—0
Task-Aware Unified Source Separation—0
Teacher-Student MixIT for Unsupervised and Semi-supervised Speech Separation—0
Temporal-Spatial Neural Filter: Direction Informed End-to-End Multi-channel Target Speech Separation—0
Tensor-Train Long Short-Term Memory for Monaural Speech Enhancement—0
The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines—0
The RoyalFlush System of Speech Recognition for M2MeT Challenge—0
TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation—0
Progressive Learning for Stabilizing Label Selection in Speech Separation with Mapping-based Method—0
Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism—0
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation—0
TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition—0
Towards Listening to 10 People Simultaneously: An Efficient Permutation Invariant Training of Audio Source Separation Using Sinkhorn's Algorithm—0
Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments—0
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition—0
Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect—0
Ultra Fast Speech Separation Model with Teacher Student Learning—0
Ultra-Lightweight Speech Separation via Group Communication—0
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation—0
UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training Mixtures—0
Unsupervised Sound Separation Using Mixture Invariant Training—0
Using Optimal Ratio Mask as Training Target for Supervised Speech Separation—0
USTC-NELSLIP System Description for DIHARD-III Challenge—0
Utterance-level Permutation Invariant Training with Latency-controlled BLSTM for Single-channel Multi-talker Speech Separation—0
VarArray: Array-Geometry-Agnostic Continuous Speech Separation—0
VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition—0
Wanna hear your voice? A sample is all we need!—0
Wavesplit: End-to-End Speech Separation by Speaker Clustering—0
X-DC: Explainable Deep Clustering based on Learnable Spectrogram Templates—0
Universal Sound Separation—0
SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate—0
Scaling strategies for on-device low-complexity source separation with Conv-Tasnet—0
SCA: Streaming Cross-attention Alignment for Echo Cancellation—0
Seeing Through the Conversation: Audio-Visual Speech Separation based on Diffusion Model—0
Self-Remixing: Unsupervised Speech Separation via Separation and Remixing—0
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation—0
Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation—0
Separating Long-Form Speech with Group-Wise Permutation Invariant Training—0
Separation Guided Speaker Diarization in Realistic Mismatched Conditions—0
Separator-Transducer-Segmenter: Streaming Recognition and Segmentation of Multi-party Speech—0
SepIt: Approaching a Single Channel Speech Separation Bound—0
Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals—0
Show:102550
← PrevPage 6 of 8Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TF-Locoformer (L) + DMSI-SDRi25.1—Unverified
2SepReformer-LSI-SDRi25.1—Unverified
3TF-Locoformer (M) + DMSI-SDRi24.6—Unverified
4TF-Locoformer (L)SI-SDRi24.2—Unverified
5MossFormer2 (L)SI-SDRi24.1—Unverified
6SepTDA (L=12)SI-SDRi24—Unverified
7Separate And DiffuseSI-SDRi23.9—Unverified
8TF-Locoformer (M)SI-SDRi23.6—Unverified
9MossFormer (L) + DMSI-SDRi22.8—Unverified
10TF-Locoformer (S) + DMSI-SDRi22.8—Unverified
#ModelMetricClaimedVerifiedStatus
1TF-Locoformer (M)SI-SDRi18.5—Unverified
2TF-Locoformer (S)SI-SDRi17.4—Unverified
3SepReformer-L + DMSI-SDRi17.1—Unverified
4MossFormer2SI-SDRi17—Unverified
5MossFormer (L) + DMSI-SDRi16.3—Unverified
6TD-Conformer (XL) + DMSI-SDRi14.6—Unverified
7Improved Sudo rm -rf (U=36)SI-SDRi13.5—Unverified
8TD-Conformer (L) + DMSI-SDRi13.4—Unverified
9WavesplitSI-SDRi13.2—Unverified
10DPTNET - SRSSNSI-SDRi12.3—Unverified
#ModelMetricClaimedVerifiedStatus
1MossFormer2 (w speed perturb)SI-SDRi22.2—Unverified
2TF-Locoformer (M)SI-SDRi22.1—Unverified
3MossFormer2 (w/o DM)SI-SDRi21.7—Unverified
4Separate And DiffuseSI-SDRi21.5—Unverified
5WHYVSI-SDRi17.5—Unverified
6TDANet LargeSI-SDRi17.4—Unverified
7TDANetSI-SDRi16.9—Unverified
8Conv-Tasnet (Libri1Mix speech enhancement pre-trained)SI-SDRi14.1—Unverified
9Conv-Tasnet (Libri1Mix speech enhancement multi-task)SI-SDRi13.7—Unverified
10Conv-TasnetSI-SDRi13.2—Unverified
#ModelMetricClaimedVerifiedStatus
1SepTDASI-SDRi23.7—Unverified
2MossFormer2SI-SDRi22.2—Unverified
3MossFormer (L) + DMSI-SDRi21.2—Unverified
4Separate And DiffuseSI-SDRi20.9—Unverified
5MossFormer (M) + DMSI-SDRi20.8—Unverified
6SepItSI-SDRi20.1—Unverified
7SepFormerSI-SDRi19.5—Unverified
8SandglassetSI-SDRi17.1—Unverified
9Gated DualPathRNNSI-SDRi16.85—Unverified
#ModelMetricClaimedVerifiedStatus
1IIANetSI-SNRi16.4—Unverified
2TDFNet-largeSI-SNRi15.8—Unverified
3TDFNet (MHSA + Shared)SI-SNRi15—Unverified
4RTFS-Net-12SI-SNRi14.9—Unverified
5RTFS-Net-6SI-SNRi14.6—Unverified
6CTCNetSI-SNRi14.3—Unverified
7RTFS-Net-4SI-SNRi14.1—Unverified
8TDFNet-smallSI-SNRi13.6—Unverified
#ModelMetricClaimedVerifiedStatus
1SepReformer-L + DMSI-SDRi18.4—Unverified
2MossFormer2SI-SDRi18.1—Unverified
3MossFormer (L) + DMSI-SDRi17.3—Unverified
4TDANet LargeSI-SDRi15.2—Unverified
5TDANetSI-SDRi14.8—Unverified
6WHYVSI-SDRi12.96—Unverified
#ModelMetricClaimedVerifiedStatus
1SepTDASI-SDRi21—Unverified
2Hungarian PITSI-SDRi13.22—Unverified
3Conditional TasNetSI-SDRi11.7—Unverified
4TasTasSI-SDRi11.14—Unverified
5Gated DualPathRNNSI-SDRi10.56—Unverified
6Multi-Decoder DPRNNSI-SDRi5.9—Unverified
#ModelMetricClaimedVerifiedStatus
1IIANetSI-SNRi18.3—Unverified
2RTFS-Net-12SI-SNRi17.5—Unverified
3CTCNetSI-SNRi17.4—Unverified
4RTFS-Net-6SI-SNRi16.9—Unverified
5RTFS-Net-4SI-SNRi15.5—Unverified
#ModelMetricClaimedVerifiedStatus
1IIANetSI-SNRi14—Unverified
2RTFS-Net-12SI-SNRi12.4—Unverified
3CTCNetSI-SNRi11.9—Unverified
4RTFS-Net-6SI-SNRi11.8—Unverified
5RTFS-Net-4SI-SNRi11.5—Unverified
#ModelMetricClaimedVerifiedStatus
1SepTDASI-SDRi22—Unverified
2Gated DualPathRNNSI-SDRi12.88—Unverified
3Conditional TasNetSI-SDRi12.5—Unverified
4OR-PITSI-SDRi10.2—Unverified
5Multi-Decoder DPRNNSI-SDRi9.3—Unverified
#ModelMetricClaimedVerifiedStatus
1Separate And DiffuseSI-SDRi14.2—Unverified
2SepItSI-SDRi13.7—Unverified
3OCDSI-SDRi13.4—Unverified
4Hungarian PITSI-SDRi12.72—Unverified
#ModelMetricClaimedVerifiedStatus
1Separate And DiffuseSI-SDRi9—Unverified
2SepItSI-SDRi8.2—Unverified
3Hungarian PITSI-SDRi7.78—Unverified
#ModelMetricClaimedVerifiedStatus
1SDR9.6—Unverified
2Audio-Visual concat-refSDR8.05—Unverified
#ModelMetricClaimedVerifiedStatus
1Separate And DiffuseSI-SDRi5.2—Unverified
2Hungarian PITSI-SDRi4.26—Unverified
#ModelMetricClaimedVerifiedStatus
1Conformer (base)0S5.6—Unverified
2Conformer (large)0S5.4—Unverified
#ModelMetricClaimedVerifiedStatus
1Hungarian PITSI-SDRi5.66—Unverified
#ModelMetricClaimedVerifiedStatus
1Audio-Visual concat-refSDR10.55—Unverified
#ModelMetricClaimedVerifiedStatus
1MossFormer2SI-SDRi20.5—Unverified