SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 851860 of 1149 papers

TitleStatusHype
Massively Parallel Video Networks0
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model0
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding0
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal Localization0
Memory Consolidation Enables Long-Context Video Understanding0
Memory-enhanced Retrieval Augmentation for Long Video Understanding0
Memory-Guided Semantic Learning Network for Temporal Sentence Grounding0
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD0
MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound0
Mid-level Representation for Visual Recognition0
Show:102550
← PrevPage 86 of 115Next →

No leaderboard results yet.