SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 691700 of 1149 papers

TitleStatusHype
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment0
EVQAScore: Efficient Video Question Answering Data Evaluation0
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding0
Egocentric and Exocentric Methods: A Short Survey0
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling0
Exploiting Spatial-Temporal Modelling and Multi-Modal Fusion for Human Action Recognition0
Exploring Anchor-based Detection for Ego4D Natural Language Query0
Exploring Missing Modality in Multimodal Egocentric Datasets0
Exploring State Change Capture of Heterogeneous Backbones @ Ego4D Hands and Objects Challenge 20220
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding0
Show:102550
← PrevPage 70 of 115Next →

No leaderboard results yet.