SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 851875 of 1149 papers

TitleStatusHype
STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition0
EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding0
PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval0
Self-Supervised Object Detection from Egocentric Videos0
Relational Space-Time Query in Long-Form Videos0
Few-Shot Referring Relationships in VideosCode0
UniFormerV2: Unlocking the Potential of Image ViTs for Video Understanding0
Inverse Compositional Learning for Weakly-supervised Relation Grounding0
Multimodal High-order Relation Transformer for Scene Boundary Detection0
Joint Engagement Classification using Video Augmentation Techniques for Multi-person Human-robot Interaction0
Inductive Attention for Video Action Anticipation0
Egocentric Video Task Translation0
Contextual Explainable Video Representation: Human Perception-based UnderstandingCode0
PromptonomyViT: Multi-Task Prompt Learning Improves Video Transformers using Synthetic Scene Data0
Transition Is a Process: Pair-to-Video Change Detection Networks for Very High Resolution Remote Sensing Images0
Spatio-Temporal Crop Aggregation for Video Representation Learning0
Dynamic Appearance: A Video Representation for Action Recognition with Joint Training0
A Unified Model for Video Understanding and Knowledge Embedding with Heterogeneous Knowledge Graph Dataset0
Masked Autoencoders for Egocentric Video Understanding @ Ego4D Challenge 2022Code0
Exploring State Change Capture of Heterogeneous Backbones @ Ego4D Hands and Objects Challenge 20220
Grounded Video Situation Recognition0
How Would The Viewer Feel? Estimating Wellbeing From Video ScenariosCode0
Self-supervised video pretraining yields robust and more human-aligned visual representations0
Students taught by multimodal teachers are superior action recognizers0
Compressed Vision for Efficient Video Understanding0
Show:102550
← PrevPage 35 of 46Next →

No leaderboard results yet.