SOTAVerified

Speech Separation

The task of extracting all overlapping speech sources in a given mixed speech signal refers to the Speech Separation. Speech Separation is a special scenario of source separation problem, where the focus is only on the overlapping speech signal sources and other interferences such as music or noise signals are not the main concern of the study. A recent representative Github project can be referred to ClearerVoice-Studio.

Source: A Unified Framework for Speech Separation

Image credit: Speech Separation of A Target Speaker Based on Deep Neural Networks

Papers

Showing 1–10 of 359 papers

TitleStatusHype
Dynamic Slimmable Networks for Efficient Speech Separation—0
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios—0
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative PipelineCode3
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers—0
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues—0
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation—0
SepPrune: Structured Pruning for Efficient Deep Speech SeparationCode1
A Survey of Deep Learning for Complex Speech Spectrograms—0
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion PriorCode1
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer—0
Show:102550
← PrevPage 1 of 36Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MossFormer2 (w speed perturb)SI-SDRi22.2—Unverified
2TF-Locoformer (M)SI-SDRi22.1—Unverified
3MossFormer2 (w/o DM)SI-SDRi21.7—Unverified
4Separate And DiffuseSI-SDRi21.5—Unverified
5WHYVSI-SDRi17.5—Unverified
6TDANet LargeSI-SDRi17.4—Unverified
7TDANetSI-SDRi16.9—Unverified
8Conv-Tasnet (Libri1Mix speech enhancement pre-trained)SI-SDRi14.1—Unverified
9Conv-Tasnet (Libri1Mix speech enhancement multi-task)SI-SDRi13.7—Unverified
10Conv-TasnetSI-SDRi13.2—Unverified