SOTAVerified|Agents Browse Leaderboard About

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 201–210 of 1149 papers

Title	Date	Tasks	Status	Hype
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding	Dec 3, 2024	In-Context LearningVideo Understanding	CodeCode Available	1
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos	Dec 2, 2024	Question AnsweringVideo Understanding	CodeCode Available	1
T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs	Nov 29, 2024	Data AugmentationDiversity	CodeCode Available	1
Teaching VLMs to Localize Specific Objects from In-context Examples	Nov 20, 2024	ObjectObject Tracking	CodeCode Available	1
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models	Nov 17, 2024	MVBenchVideo-based Generative Performance Benchmarking	CodeCode Available	1
Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation	Oct 31, 2024	Action SegmentationAction Understanding	CodeCode Available	1
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models	Oct 30, 2024	Video Understanding	CodeCode Available	1
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks	Oct 24, 2024	Video Understanding	CodeCode Available	1
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark	Oct 24, 2024	document understandingVideo Understanding	CodeCode Available	1
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models	Oct 14, 2024	2kBenchmarking	CodeCode Available	1

Show:10 25 50

← PrevPage 21 of 115Next →

No leaderboard results yet.