SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 401410 of 1149 papers

TitleStatusHype
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks0
EAGLE: Egocentric AGgregated Language-video Engine0
LLM4Brain: Training a Large Language Model for Brain Video Understanding0
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingCode2
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP0
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video UnderstandingCode4
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge0
Towards Child-Inclusive Clinical Video Understanding for Autism Spectrum Disorder0
Interpretable Action Recognition on Hard to Classify Actions0
AMEGO: Active Memory from long EGOcentric videos0
Show:102550
← PrevPage 41 of 115Next →

No leaderboard results yet.