SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 371380 of 1149 papers

TitleStatusHype
EVA: An Embodied World Model for Future Video Anticipation0
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning0
Making Every Frame Matter: Continuous Video Understanding for Large Models via Adaptive State Modeling0
Zero-shot Action Localization via the Confidence of Large Vision-Language Models0
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AICode2
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models0
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video ModelsCode1
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMsCode2
ViFi-ReID: A Two-Stream Vision-WiFi Multimodal Approach for Person Re-identification0
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering0
Show:102550
← PrevPage 38 of 115Next →

No leaderboard results yet.