SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 451475 of 1149 papers

TitleStatusHype
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles0
ViQAgent: Zero-Shot Video Question Answering via Agent with Open-Vocabulary Grounding ValidationCode0
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval0
Clapper: Compact Learning and Video Representation in VLMs0
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning0
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition0
Domain Adaptation of VLM for Soccer Video Understanding0
A Challenge to Build Neuro-Symbolic Video AgentsCode0
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?0
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video UnderstandingCode0
From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations0
SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation0
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language ModelsCode0
Gameplay Highlights Generation0
Seed1.5-VL Technical Report0
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant0
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph0
VideoLLM Benchmarks and Evaluation: A Survey0
Empowering Agentic Video Analytics Systems with Video Language Models0
SeriesBench: A Benchmark for Narrative-Driven Drama Series UnderstandingCode0
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation0
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs0
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes0
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding0
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection0
Show:102550
← PrevPage 19 of 46Next →

No leaderboard results yet.