SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 321330 of 1149 papers

TitleStatusHype
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model0
VisionZip: Longer is Better but Not Necessary in Vision Language ModelsCode3
Streaming Detection of Queried Event StartCode0
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningCode2
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding0
Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction TuningCode1
Progress-Aware Video Frame Captioning0
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video UnderstandingCode1
PhysGame: Uncovering Physical Commonsense Violations in Gameplay VideosCode1
SEAL: Semantic Attention Learning for Long Video Representation0
Show:102550
← PrevPage 33 of 115Next →

No leaderboard results yet.