SOTAVerified

Action Understanding

Papers

Showing 1–50 of 88 papers

TitleStatusHype
OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow UnderstandingCode2
LLaVAction: evaluating and training multi-modal large language models for action recognitionCode2
Bridge-Prompt: Towards Ordinal Action Understanding in Instructional VideosCode1
Memory-and-Anticipation Transformer for Online Action UnderstandingCode1
Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic ActionsCode1
Action Quality Assessment with Temporal Parsing TransformerCode1
YouMakeup VQA Challenge: Towards Fine-grained Action Understanding in Domain-Specific VideosCode1
Towards Tokenized Human Dynamics RepresentationCode1
Open-Vocabulary Video Relation ExtractionCode1
Paxion: Patching Action Knowledge in Video-Language Foundation ModelsCode1
Home Action Genome: Cooperative Compositional Action UnderstandingCode1
PIANO: A Parametric Hand Bone Model from Magnetic Resonance ImagingCode1
F^3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from VideosCode1
Video Pose Distillation for Few-Shot, Fine-Grained Sports Action RecognitionCode1
Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningCode1
Detailed 2D-3D Joint Representation for Human-Object InteractionCode1
FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality AssessmentCode1
FineSports: A Multi-person Hierarchical Sports Video Dataset for Fine-grained Action UnderstandingCode1
Domain Knowledge-Informed Self-Supervised Representations for Workout Form AssessmentCode1
SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationCode1
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action UnderstandingCode1
Temporal Relational Modeling with Self-Supervision for Action SegmentationCode1
Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action SegmentationCode1
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional SportsCode1
LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task ActivitiesCode1
Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingCode1
Human Action Segmentation With Hierarchical Supervoxel Consistency—0
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding—0
Impact of Large Language Model Assistance on Patients Reading Clinical Notes: A Mixed-Methods Study—0
Intra- and Inter-Action Understanding via Temporal Action Parsing—0
Invisible-to-Visible: Privacy-Aware Human Instance Segmentation using Airborne Ultrasound via Collaborative Learning Variational Autoencoder—0
JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection—0
Kantian Deontology Meets AI Alignment: Towards Morally Grounded Fairness Metrics—0
MacDiff: Unified Skeleton Modeling with Masked Conditional Diffusion—0
MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding—0
mRI: Multi-modal 3D Human Pose Estimation Dataset using mmWave, RGB-D, and Inertial Sensors—0
Multitask Learning in Minimally Invasive Surgical Vision: A Review—0
PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition—0
PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding—0
Probing Fine-Grained Action Understanding and Cross-View Generalization of Foundation Models—0
Region-aware Image-based Human Action Retrieval with Transformers—0
RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics—0
Scene Understanding for Autonomous Manipulation with Deep Learning—0
ScreenLLM: Stateful Screen Schema for Efficient Action Understanding and Prediction—0
Self-supervised Discovery of Human Actons from Long Kinematic Videos—0
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning—0
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding—0
The SkatingVerse Workshop & Challenge: Methods and Results—0
Action Understanding with Multiple Classes of Actors—0
Actor and Action Modular Network for Text-based Video Segmentation—0
Show:102550
← PrevPage 1 of 2Next →

No leaderboard results yet.