SOTAVerified

Activity Detection

Detecting activities in extended videos.

Papers

Showing 251300 of 380 papers

TitleStatusHype
Tackling the Cocktail Fork Problem for Separation and Transcription of Real-World Soundtracks0
Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription0
Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario0
Target-Speaker Voice Activity Detection via Sequence-to-Sequence Prediction0
Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker0
Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization0
TCG CREST System Description for the Second DISPLACE Challenge0
Spatio-Temporal Event Segmentation and Localization for Wildlife Extended Videos0
Temporarily-Aware Context Modelling using Generative Adversarial Networks for Speech Activity Detection0
Tensor vs Matrix Methods: Robust Tensor Decomposition under Block Sparse Perturbations0
The AFRL IWSLT 2020 Systems: Work-From-Home Edition0
The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge0
The DKU-DukeECE Diarization System for the VoxCeleb Speaker Recognition Challenge 20220
The DKU-DukeECE-Lenovo System for the Diarization Task of the 2021 VoxCeleb Speaker Recognition Challenge0
The DKU-MSXF Diarization System for the VoxCeleb Speaker Recognition Challenge 20230
The HUAWEI Speaker Diarisation System for the VoxCeleb Speaker Diarisation Challenge0
The Impact of Silence on Speech Anti-Spoofing0
The JHU Multi-Microphone Multi-Speaker ASR System for the CHiME-6 Challenge0
The Kriston AI System for the VoxCeleb Speaker Recognition Challenge 20220
The Newsbridge -Telecom SudParis VoxCeleb Speaker Recognition Challenge 2022 System Description0
The RATS Collection: Supporting HLT Research with Degraded Audio Data0
The SAFE-T Corpus: A New Resource for Simulated Public Safety Communications0
The "Sound of Silence" in EEG -- Cognitive voice activity detection0
The Speed Submission to DIHARD II: Contributions & Lessons Learned0
The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge0
The VVAD-LRS3 Dataset for Visual Voice Activity Detection0
"This is Houston. Say again, please". The Behavox system for the Apollo-11 Fearless Steps Challenge (phase II)0
Token Turing Machines0
Towards end-2-end learning for predicting behavior codes from spoken utterances in psychotherapy conversations0
Towards More Practical Group Activity Detection: A New Benchmark and Model0
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM0
Trajectory-User Linking Is Easier Than You Think0
Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR0
Transferable Adversarial Attacks against ASR0
TRECVID 2019: An Evaluation Campaign to Benchmark Video Activity Detection, Video Captioning and Matching, and Video Search & Retrieval0
Tri-axial Self-Attention for Concurrent Activity Recognition0
TSUP Speaker Diarization System for Conversational Short-phrase Speaker Diarization Challenge0
Two-stream Multi-dimensional Convolutional Network for Real-time Violence Detection0
Two-Stream Region Convolutional 3D Network for Temporal Activity Detection0
Ultra-sensitive Flexible Sponge-Sensor Array for Muscle Activities Detection and Human Limb Motion Recognition0
Union of Low-Rank Subspaces Detector0
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection0
Unveiling ECC Vulnerabilities: LSTM Networks for Operation Recognition in Side-Channel Attacks0
Unveiling the Power of Complex-Valued Transformers in Wireless Communications0
User Activity Detection and Channel Estimation of Spatially Correlated Channels via AMP in Massive MTC0
User Activity Detection for Irregular Repetition Slotted Aloha based MMTC0
User Activity Detection with Delay-Calibration for Asynchronous Massive Random Access0
USTC-NELSLIP System Description for DIHARD-III Challenge0
VAD-free Streaming Hybrid CTC/Attention ASR for Unsegmented Recording0
VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition0
Show:102550
← PrevPage 6 of 8Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1CNN-BiLSTM_bestROC-AUC95.14Unverified
2CNN-BiLSTM_smallROC-AUC95.13Unverified
3SG-VAD (ours)ROC-AUC94.3Unverified
4ADA-VADROC-AUC79.1Unverified