SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 476500 of 1149 papers

TitleStatusHype
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task0
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding0
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?0
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval0
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization0
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild0
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding0
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model0
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking0
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding0
How Can Objects Help Video-Language Understanding?0
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding0
From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction0
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models0
InstructionBench: An Instructional Video Understanding Benchmark0
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval0
Moment Quantization for Video Temporal Grounding0
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding0
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?0
Aligned Better, Listen Better for Audio-Visual Large Language Models0
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description0
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding0
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition0
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts0
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding0
Show:102550
← PrevPage 20 of 46Next →

No leaderboard results yet.