SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 211220 of 1149 papers

TitleStatusHype
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video UnderstandingCode1
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMsCode1
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video UnderstandingCode1
HAT: History-Augmented Anchor Transformer for Online Temporal Action LocalizationCode1
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language BenchmarkCode1
Learning Video Context as Interleaved Multimodal SequencesCode1
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video RetrievalCode1
Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video UnderstandingCode1
VideoMamba: Spatio-Temporal Selective State Space ModelCode1
MMAD: Multi-label Micro-Action Detection in VideosCode1
Show:102550
← PrevPage 22 of 115Next →

No leaderboard results yet.