SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 781790 of 1149 papers

TitleStatusHype
Dr^2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient FinetuningCode0
VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding0
Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning0
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action0
Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning0
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding0
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision0
Text-Conditioned Resampler For Long Form Video Understanding0
Learning Object State Changes in Videos: An Open-World Perspective0
Artificial intelligence optical hardware empowers high-resolution hyperspectral video understanding at 1.2 Tb/s0
Show:102550
← PrevPage 79 of 115Next →

No leaderboard results yet.