SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 876900 of 1149 papers

TitleStatusHype
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning0
MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding0
MRSN: Multi-Relation Support Network for Video Action Detection0
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language0
Multi-kernel learning of deep convolutional features for action recognition0
Multimodal High-order Relation Transformer for Scene Boundary Detection0
Multimodal Intent Discovery from Livestream Videos0
Multi-modal Representation Learning for Video Advertisement Content Structuring0
Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation0
Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding0
Multi-scale 2D Temporal Map Diffusion Models for Natural Language Video Localization0
Multi-Scale Contrastive Learning for Video Temporal Grounding0
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding0
Multiview Transformers for Video Recognition0
MVTamperBench: Evaluating Robustness of Vision-Language Models0
Representation Learning on Visual-Symbolic Graphs for Video Understanding0
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision0
Non-local NetVLAD Encoding for Video Classification0
O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning0
OBJECT DYNAMICS DISTILLATION FOR SCENE DECOMPOSITION AND REPRESENTATION0
Occluded Video Instance Segmentation: Dataset and ICCV 2021 Challenge0
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding0
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts0
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks0
OmniTrack: Real-time detection and tracking of objects, text and logos in video0
Show:102550
← PrevPage 36 of 46Next →

No leaderboard results yet.