Action Detection

Action Detection aims to find both where and when an action occurs within a video clip and classify what the action is taking place. Typically results are given in the form of action tublets, which are action bounding boxes linked across time in the video. This is related to temporal localization, which seeks to identify the start and end frame of an action, and action recognition, which seeks only to classify which action is taking place and typically assumes a trimmed video.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–75 of 817 papers

Title	Date	Tasks	Status	Hype
Post-Processing Temporal Action Detection	Nov 27, 2022	Action ClassificationAction Detection	CodeCode Available	1
Multi-Modal Few-Shot Temporal Action Detection	Nov 27, 2022	Action DetectionFew-Shot Object Detection	CodeCode Available	1
Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization	Nov 12, 2022	Action DetectionActivity Detection	CodeCode Available	1
SG-VAD: Stochastic Gates Based Speech Activity Detection	Oct 28, 2022	Action DetectionActivity Detection	CodeCode Available	1
Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0	Oct 26, 2022	Action DetectionActivity Detection	CodeCode Available	1
Brouhaha: multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimation	Oct 24, 2022	Action DetectionActivity Detection	CodeCode Available	1
Holistic Interaction Transformer Network for Action Detection	Oct 23, 2022	Action DetectionAction Recognition	CodeCode Available	1
PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points	Oct 20, 2022	Action DetectionTemporal Action Localization	CodeCode Available	1
YOWO-Plus: An Incremental Improvement	Oct 20, 2022	Action DetectionGPU	CodeCode Available	1
AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation	Oct 5, 2022	Action DetectionTemporal Action Proposal Generation	CodeCode Available	1
Exploiting Instance-based Mixed Sampling via Auxiliary Source Domain Supervision for Domain-adaptive Action Detection	Sep 28, 2022	Action DetectionDomain Adaptation	CodeCode Available	1
Real-time Online Video Detection with Temporal Smoothing Transformers	Sep 19, 2022	Action AnticipationAction Detection	CodeCode Available	1
Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions	Jul 24, 2022	Action DetectionAction Understanding	CodeCode Available	1
Spotting Temporally Precise, Fine-Grained Events in Video	Jul 20, 2022	Action DetectionAction Spotting	CodeCode Available	1
Hierarchically Self-Supervised Transformer for Human Skeleton Representation Learning	Jul 20, 2022	Action DetectionAction Recognition	CodeCode Available	1
Zero-Shot Temporal Action Detection via Vision-Language Prompting	Jul 17, 2022	Action DetectionClassification	CodeCode Available	1
Proposal-Free Temporal Action Detection via Global Segmentation Mask Learning	Jul 14, 2022	Action DetectionRepresentation Learning	CodeCode Available	1
Semi-Supervised Temporal Action Detection with Proposal-Free Masking	Jul 14, 2022	Action DetectionGeneral Classification	CodeCode Available	1
ReAct: Temporal Action Detection with Relational Queries	Jul 14, 2022	Action ClassificationAction Detection	CodeCode Available	1
MM-ALT: A Multimodal Automatic Lyric Transcription System	Jul 13, 2022	Action DetectionActivity Detection	CodeCode Available	1
A semi-supervised methodology for fishing activity detection using the geometry behind the trajectory of multiple vessels	Jul 12, 2022	Action DetectionActivity Detection	CodeCode Available	1
Unsupervised Voice Activity Detection by Modeling Source and System Information using Zero Frequency Filtering	Jun 27, 2022	Action DetectionActivity Detection	CodeCode Available	1
A Simple and Efficient Pipeline to Build an End-to-End Spatial-Temporal Action Detector	Jun 7, 2022	Action ClassificationAction Detection	CodeCode Available	1
Stargazer: A transformer-based driver action detection system for intelligent transportation	Jun 1, 2022	Action DetectionAction Recognition	CodeCode Available	1
ETAD: Training Action Detection End to End on a Laptop	May 14, 2022	Action DetectionGPU	CodeCode Available	1

Show:10 25 50

← PrevPage 3 of 33Next →

All datasets UCF101-24 J-HMDB Charades Multi-THUMOS UCF Sports THUMOS' 14 MultiSports TSU TTStroke-21 ME21 TTStroke-21 ME22 MultiTHUMOS

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	STAR/L	Frame-mAP 0.5	90.3	—	Unverified
2	SiA	Frame-mAP 0.5	88.5	—	Unverified
3	YOWO + LFB	Frame-mAP 0.5	87.3	—	Unverified
4	HIT	Frame-mAP 0.5	84.8	—	Unverified
5	HISAN (ResNet-101 + FPN)	Video-mAP 0.2	82.3	—	Unverified
6	YOWO	Frame-mAP 0.5	80.4	—	Unverified
7	Two-in-one Two Stream	Video-mAP 0.2	78.48	—	Unverified
8	MOC	Frame-mAP 0.5	77.8	—	Unverified
9	Faster-RCNN + two-stream I3D conv	Frame-mAP 0.5	76.3	—	Unverified
10	Two-in-one	Video-mAP 0.2	75.48	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	SiA	Frame-mAP 0.5	88.5	—	Unverified
2	HISAN (ResNet-101 + FPN)	Video-mAP 0.2	87.59	—	Unverified
3	HIT	Frame-mAP 0.5	83.8	—	Unverified
4	HISAN (VGG-16)	Frame-mAP 0.5	76.72	—	Unverified
5	DTS	Video-mAP 0.2	76.1	—	Unverified
6	YOWO + LFB	Frame-mAP 0.5	75.7	—	Unverified
7	Two-in-one Two Stream	Video-mAP 0.5	74.74	—	Unverified
8	YOWO	Frame-mAP 0.5	74.4	—	Unverified
9	MOC	Frame-mAP 0.5	74	—	Unverified
10	Faster-RCNN + two-stream I3D conv	Frame-mAP 0.5	73.3	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	TTM	mAP	28.79	—	Unverified
2	CTRN	mAP	27.8	—	Unverified
3	Coarse-Fine Networks (w/ self-supervised detection pretraining)	mAP	26.95	—	Unverified
4	UniMD+Sync. (RGB+Flow)	mAP	26.53	—	Unverified
5	PDAN (RGB+Flow)	mAP	26.5	—	Unverified
6	PAT	mAP	26.5	—	Unverified
7	MS-TCT (RGB only)	mAP	25.4	—	Unverified
8	3D ResNet-50 + super-events pretrained on AViD	mAP	25.2	—	Unverified
9	Coarse-Fine Networks	mAP	25.1	—	Unverified
10	I3D + biGRU + VS-ST-MPNN	mAP	23.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	MLAD	mAP	51.5	—	Unverified
2	CTRN	mAP	51.2	—	Unverified
3	PDAN	mAP	47.6	—	Unverified
4	TGM	mAP	46.4	—	Unverified
5	MS-TCT (RGB only)	mAP	43.1	—	Unverified
6	I3D + our super-event	mAP	36.4	—	Unverified
7	Two-stream + LSTM	mAP	28.1	—	Unverified
8	Two-stream	mAP	27.6	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	Two-in-one Two Stream	Video-mAP 0.5	96.52	—	Unverified
2	DTS	Video-mAP 0.2	94.3	—	Unverified
3	Two-in-one	Video-mAP 0.5	92.74	—	Unverified
4	T-CNN	Frame-mAP 0.5	86.7	—	Unverified
5	MR-TS R-CNN	Frame-mAP 0.5	84.52	—	Unverified
6	TS R-CNN	Frame-mAP 0.5	82.3	—	Unverified
7	Action Tubes	Frame-mAP 0.5	68.1	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	MAT (Ours) Trans	mAP	71.6	—	Unverified
2	TadML-two stream	mAP	59.7	—	Unverified
3	MAT (ours)	mAP	58.2	—	Unverified
4	TadML-rgb	mAP	53.46	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	HIT	Frame-mAP 0.5	33.3	—	Unverified
2	SiA	Frame-mAP 0.5	28.8	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	MS-TCT	Frame-mAP	33.7	—	Unverified
2	PDAN	Frame-mAP	32.7	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	STCNN	IoU	0.14	—	Unverified
2	Two Stream Network	IoU	0.07	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	STCNN-V2 (Vote decision)	IoU	0.52	—	Unverified
2	RGB and PRGB	IoU	0.35	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	PAT	mAP	44.6	—	Unverified