SOTAVerified

Action Understanding

Papers

Showing 125 of 88 papers

TitleStatusHype
LLaVAction: evaluating and training multi-modal large language models for action recognitionCode2
OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow UnderstandingCode2
F^3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from VideosCode1
SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationCode1
Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action SegmentationCode1
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action UnderstandingCode1
FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality AssessmentCode1
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional SportsCode1
FineSports: A Multi-person Hierarchical Sports Video Dataset for Fine-grained Action UnderstandingCode1
Open-Vocabulary Video Relation ExtractionCode1
Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingCode1
Memory-and-Anticipation Transformer for Online Action UnderstandingCode1
Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningCode1
Paxion: Patching Action Knowledge in Video-Language Foundation ModelsCode1
Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic ActionsCode1
Action Quality Assessment with Temporal Parsing TransformerCode1
Bridge-Prompt: Towards Ordinal Action Understanding in Instructional VideosCode1
Domain Knowledge-Informed Self-Supervised Representations for Workout Form AssessmentCode1
Towards Tokenized Human Dynamics RepresentationCode1
Video Pose Distillation for Few-Shot, Fine-Grained Sports Action RecognitionCode1
PIANO: A Parametric Hand Bone Model from Magnetic Resonance ImagingCode1
Home Action Genome: Cooperative Compositional Action UnderstandingCode1
Temporal Relational Modeling with Self-Supervision for Action SegmentationCode1
LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task ActivitiesCode1
Detailed 2D-3D Joint Representation for Human-Object InteractionCode1
Show:102550
← PrevPage 1 of 4Next →

No leaderboard results yet.