SOTAVerified

Action Detection

Action Detection aims to find both where and when an action occurs within a video clip and classify what the action is taking place. Typically results are given in the form of action tublets, which are action bounding boxes linked across time in the video. This is related to temporal localization, which seeks to identify the start and end frame of an action, and action recognition, which seeks only to classify which action is taking place and typically assumes a trimmed video.

Papers

Showing 1–10 of 817 papers

TitleStatusHype
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans—0
CBF-AFA: Chunk-Based Multi-SSL Fusion for Automatic Fluency Assessment—0
Distributed Activity Detection for Cell-Free Hybrid Near-Far Field Communications—0
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation AlgorithmCode1
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion—0
Joint Activity Detection and Channel Estimation for Massive Connectivity: Where Message Passing Meets Score-Based Generative Priors—0
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM—0
Robust Activity Detection for Massive Random Access—0
Improving endpoint detection in end-to-end streaming ASR for conversational speech—0
Multi-Stage Speaker Diarization for Noisy ClassroomsCode0
Show:102550
← PrevPage 1 of 82Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1MAT (Ours) TransmAP71.6—Unverified
2TadML-two streammAP59.7—Unverified
3MAT (ours)mAP58.2—Unverified
4TadML-rgbmAP53.46—Unverified