SOTAVerified

Temporal Action Localization

Temporal Action Localization aims to detect activities in the video stream and output beginning and end timestamps. It is closely related to Temporal Action Proposal Generation.

Papers

Showing 1–50 of 1477 papers

TitleStatusHype
InternVideo2: Scaling Foundation Models for Multimodal Video UnderstandingCode7
InternVideo: General Video Foundation Models via Generative and Discriminative LearningCode4
Video Mamba Suite: State Space Model as a Versatile Alternative for Video UnderstandingCode3
A Survey on Video Action Recognition in Sports: Datasets, Methods and ApplicationsCode3
Structured Attention Composition for Temporal Action LocalizationCode2
Leveraging Temporal Contextualization for Video Action RecognitionCode2
AIM: Adapting Image Models for Efficient Video Action RecognitionCode2
ActionFormer: Localizing Moments of Actions with TransformersCode2
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationCode2
TriDet: Temporal Action Detection with Relative Boundary ModelingCode2
Test-Time Zero-Shot Temporal Action LocalizationCode2
NMS Threshold matters for Ego4D Moment Queries -- 2nd place solution to the Ego4D Moment Queries Challenge 2023Code2
End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesCode2
Perception Test: A Diagnostic Benchmark for Multimodal Video ModelsCode2
The Surprising Effectiveness of Multimodal Large Language Models for Video Moment RetrievalCode2
On the Benefits of 3D Pose and Tracking for Human Action RecognitionCode2
Temporal Action Localization with Enhanced Instant DiscriminabilityCode2
Temporal Segment Networks: Towards Good Practices for Deep Action RecognitionCode2
UniMD: Towards Unifying Moment Retrieval and Temporal Action DetectionCode2
Where a Strong Backbone Meets Strong Features -- ActionFormer for Ego4D Moment Queries ChallengeCode2
Temporal Segment Networks for Action Recognition in VideosCode2
VideoMAE V2: Scaling Video Masked Autoencoders with Dual MaskingCode2
Cross-modal Consensus Network forWeakly Supervised Temporal Action LocalizationCode1
Convex Combination Consistency between Neighbors for Weakly-supervised Action LocalizationCode1
CZU-MHAD: A multimodal dataset for human action recognition utilizing a depth camera and 10 wearable inertial sensorsCode1
Complex Sequential Understanding through the Awareness of Spatial and Temporal ConceptsCode1
A Closer Look at Spatiotemporal Convolutions for Action RecognitionCode1
ACM-Net: Action Context Modeling Network for Weakly-Supervised Temporal Action LocalizationCode1
Co-occurrence Feature Learning from Skeleton Data for Action Recognition and Detection with Hierarchical AggregationCode1
Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationCode1
CoLA: Weakly-Supervised Temporal Action Localization with Snippet Contrastive LearningCode1
Compressing Recurrent Neural Networks with Tensor Ring for Action RecognitionCode1
CDFSL-V: Cross-Domain Few-Shot Learning for VideosCode1
CAKES: Channel-wise Automatic KErnel Shrinking for Efficient 3D NetworksCode1
Challenges in Video-Based Infant Action Recognition: A Critical Examination of the State of the ArtCode1
Bottom-Up Temporal Action Localization with Mutual RegularizationCode1
DCAN: Improving Temporal Action Detection via Dual Context AggregationCode1
Boosting Weakly-Supervised Temporal Action Localization with Text InformationCode1
Background Suppression Network for Weakly-supervised Temporal Action LocalizationCode1
Boundary-sensitive Pre-training for Temporal Localization in VideosCode1
BABEL: Bodies, Action and Behavior with English LabelsCode1
B2C-AFM: Bi-Directional Co-Temporal and Cross-Spatial Attention Fusion Model for Human Action RecognitionCode1
ActionCLIP: A New Paradigm for Video Action RecognitionCode1
ACTION-Net: Multipath Excitation for Action RecognitionCode1
BasicTAD: an Astounding RGB-Only Baseline for Temporal Action DetectionCode1
BMN: Boundary-Matching Network for Temporal Action Proposal GenerationCode1
Background-Click Supervision for Temporal Action LocalizationCode1
BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal GenerationCode1
CBR-Net: Cascade Boundary Refinement Network for Action Detection: Submission to ActivityNet Challenge 2020 (Task 1)Code1
AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual ActionsCode1
Show:102550
← PrevPage 1 of 30Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1AdaTAD (VideoMAEv2-giant)Avg mAP (0.3:0.7)76.9—Unverified
2RDFA-S6 (InternVideo2-6B)Avg mAP (0.3:0.7)74.2—Unverified
3ActionMamba(InternVideo2-6B)Avg mAP (0.3:0.7)72.72—Unverified
4GCMmAP [email protected]72.5—Unverified
5AGT (Ours)mAP [email protected]72.1—Unverified
6InternVideo2-6BAvg mAP (0.3:0.7)72—Unverified
7ActionFormer (InternVideo features)Avg mAP (0.3:0.7)71.58—Unverified
8TriDet (VideoMAE v2-g feature)Avg mAP (0.3:0.7)70.1—Unverified
9InternVideo2-1BAvg mAP (0.3:0.7)69.8—Unverified
10ActionFormer (VideoMAE V2-g features)Avg mAP (0.3:0.7)69.6—Unverified
#ModelMetricClaimedVerifiedStatus
1UnLoc-LmAP [email protected]59.3—Unverified
2RDFA-S6 (InternVideo2-6B)mAP42.9—Unverified
3ActionMamba (InternVideo2-6B)mAP42.02—Unverified
4PRN+BMN (ensemble)mAP42—Unverified
5AdaTAD (VideoMAEv2-giant)mAP41.93—Unverified
6InternVideo2-6BmAP41.2—Unverified
7InternVideo2-1BmAP40.4—Unverified
8UniMD+Sync.mAP39.83—Unverified
9PRN (CSN)mAP39.4—Unverified
10InternVideomAP39—Unverified
#ModelMetricClaimedVerifiedStatus
1RDFA-S6 (InternVideo2-6B)Average-mAP45.8—Unverified
2ActionMamba(InternVideo2-6B)Average-mAP44.56—Unverified
3DyFADet(VideoMAEv2)Average-mAP44.3—Unverified
4InternVideo2-6BAverage-mAP43.3—Unverified
5TriDet (VideoMAEv2)Average-mAP43.1—Unverified
6InternVideo2-1BAverage-mAP42.4—Unverified
7InternVideoAverage-mAP41.55—Unverified
8TriDet (SlowFast)Average-mAP38.6—Unverified
9TriDet (I3D RGB)Average-mAP36.8—Unverified
10TadTr (I3D RGB)Average-mAP32.09—Unverified
#ModelMetricClaimedVerifiedStatus
1RDFA-S6 (InternVideo2-6B)mAP29.6—Unverified
2ActionMamba(InternVideo2-6B)mAP29.04—Unverified
3InternVideo2-6BmAP27.7—Unverified
4DyFADet (VideoMAE v2-g)mAP23.8—Unverified
5VideoMAE V2-gmAP18.24—Unverified
6InternVideomAP17.57—Unverified
7BMN (i3d feaure)mAP9.25—Unverified
8G-TAD (i3d feature)mAP9.06—Unverified
9DBG (i3d feature)mAP6.75—Unverified
#ModelMetricClaimedVerifiedStatus
1TriDet (VideoMAEv2)Average mAP37.5—Unverified
2DualDETR (I3D-rgb)Average mAP32.64—Unverified
3TriDet (I3D-rgb)Average mAP30.7—Unverified
4TemporalMaxerAverage mAP29.9—Unverified
5PointTADAverage mAP23.5—Unverified
6PDANAverage mAP17.3—Unverified
7MS-TCTAverage mAP16.2—Unverified
8MLADAverage mAP14.2—Unverified
#ModelMetricClaimedVerifiedStatus
1VideoCLIPRecall47.3—Unverified
2VLMRecall46.5—Unverified
3TACoRecall42.5—Unverified
4Text-Video EmbeddingRecall33.6—Unverified
5Fully-supervised upper-boundRecall31.6—Unverified
6ZhukovRecall22.4—Unverified
7AlayracRecall13.3—Unverified
#ModelMetricClaimedVerifiedStatus
1AdaTAD (verb, VideoMAE-L)Avg mAP (0.1-0.5)29.3—Unverified
2TriDet (verb)Avg mAP (0.1-0.5)25.4—Unverified
3TemporalMaxer (verb)Avg mAP (0.1-0.5)24.5—Unverified
4ActionFormer (verb)Avg mAP (0.1-0.5)23.5—Unverified
5G-TAD (verb)Avg mAP (0.1-0.5)9.4—Unverified
6BMN (verb)Avg mAP (0.1-0.5)8.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TemporalMaxermAP27.2—Unverified
2MUSESmAP18.6—Unverified
#ModelMetricClaimedVerifiedStatus
1DeepMetricLearnermAP [email protected]35.2—Unverified
#ModelMetricClaimedVerifiedStatus
1ActionFormer (SlowFast+Omnivore+EgoVLP)Average mAP21.76—Unverified
#ModelMetricClaimedVerifiedStatus
1ActionFormer (SlowFast+Omnivore+EgoVLP)Average mAP21.4—Unverified
#ModelMetricClaimedVerifiedStatus
1S-CNNmAP7.4—Unverified