SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 10011010 of 1149 papers

TitleStatusHype
O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning0
OBJECT DYNAMICS DISTILLATION FOR SCENE DECOMPOSITION AND REPRESENTATION0
Occluded Video Instance Segmentation: Dataset and ICCV 2021 Challenge0
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding0
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts0
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks0
OmniTrack: Real-time detection and tracking of objects, text and logos in video0
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding0
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding0
Only Time Can Tell: Discovering Temporal Data for Temporal Modeling0
Show:102550
← PrevPage 101 of 115Next →

No leaderboard results yet.