SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 931940 of 1149 papers

TitleStatusHype
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval0
LiVLR: A Lightweight Visual-Linguistic Reasoning Framework for Video Question Answering0
LLaVA-MLB: Mitigating and Leveraging Attention Bias for Training-Free Video LLMs0
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding0
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living0
LLM4Brain: Training a Large Language Model for Brain Video Understanding0
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs0
Localizing Events in Videos with Multimodal Queries0
Localizing Unseen Activities in Video via Image Query0
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding0
Show:102550
← PrevPage 94 of 115Next →

No leaderboard results yet.