SOTAVerified

Target Speaker Extraction

Extract the dialogue content of the specified target in a multi-person dialogue.

Papers

Showing 1–50 of 55 papers

TitleStatusHype
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction—0
M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker ExtractionCode0
FlowTSE: Target Speaker Extraction with Flow Matching—0
Listen to Extract: Onset-Prompted Target Speaker Extraction—0
LauraTSE: Target Speaker Extraction using Auto-Regressive Decoder-Only Language ModelsCode1
C^2AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction—0
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments—0
Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingCode9
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement—0
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection—0
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues—0
Multi-Level Speaker Representation for Target Speaker ExtractionCode3
STCON System for the CHiME-8 Challenge—0
Wanna hear your voice? A sample is all we need!—0
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions—0
WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker ExtractionCode3
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration—0
TSELM: Target Speaker Extraction using Discrete Tokens and Language ModelsCode2
USEF-TSE: Universal Speaker Embedding Free Target Speaker ExtractionCode1
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial RefinementCode0
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning—0
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling—0
Binaural Selective Attention Model for Target Speaker Extraction—0
AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band ModelingCode1
Target Speaker Extraction with Curriculum Learning—0
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory AttentionCode1
Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training—0
Target Speaker Extraction by Directly Exploiting Contextual Information in the Time-Frequency Domain—0
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives—0
A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction—0
Typing to Listen at the Cocktail Party: Text-Guided Target Speaker ExtractionCode1
Conditional Diffusion Model for Target Speaker Extraction—0
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationCode1
The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction—0
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer—0
Beamformer-Guided Target Speaker Extraction—0
Multi-Channel Target Speaker Extraction with Refinement: The WavLab Submission to the Second Clarity Enhancement Challenge—0
Improving Target Speaker Extraction with Sparse LDA-transformed Speaker Embeddings—0
GPU-accelerated Guided Source Separation for Meeting TranscriptionCode1
ExARN: self-attending RNN for target speaker extraction—0
Adapting self-supervised models to multi-talker speech recognition using speaker embeddings—0
ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding InpaintingCode0
Exploiting spatial information with the informed complex-valued spatial autoencoder for target speaker extraction—0
Semi-supervised Time Domain Target Speaker Extraction with Attention—0
Speaker-conditioning Single-channel Target Speaker Extraction using Conformer-based Architectures—0
A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker ExtractionCode1
Coarse-to-Fine Recursive Speech Separation for Unknown Number of Speakers—0
L-SpEx: Localized Target Speaker ExtractionCode1
New Insights on Target Speaker Extraction—0
Selective Listening by Synchronizing Speech with LipsCode1
Show:102550
← PrevPage 1 of 2Next →

No leaderboard results yet.