SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 831840 of 1149 papers

TitleStatusHype
Localizing Events in Videos with Multimodal Queries0
Localizing Unseen Activities in Video via Image Query0
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding0
Long Activity Video Understanding using Functional Object-Oriented Network0
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models0
Long-Short Temporal Contrastive Learning of Video Transformers0
LongVILA: Scaling Long-Context Visual Language Models for Long Videos0
LongViTU: Instruction Tuning for Long-Form Video Understanding0
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory0
Look Every Frame All at Once: Video-Ma^2mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing0
Show:102550
← PrevPage 84 of 115Next →

No leaderboard results yet.