SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 201210 of 1149 papers

TitleStatusHype
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video UnderstandingCode1
PhysGame: Uncovering Physical Commonsense Violations in Gameplay VideosCode1
T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMsCode1
Teaching VLMs to Localize Specific Objects from In-context ExamplesCode1
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language ModelsCode1
Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action SegmentationCode1
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation ModelsCode1
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web TasksCode1
CAMEL-Bench: A Comprehensive Arabic LMM BenchmarkCode1
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video ModelsCode1
Show:102550
← PrevPage 21 of 115Next →

No leaderboard results yet.