SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 451–500 of 1149 papers

TitleStatusHype
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles—0
ViQAgent: Zero-Shot Video Question Answering via Agent with Open-Vocabulary Grounding ValidationCode0
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition—0
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning—0
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval—0
Clapper: Compact Learning and Video Representation in VLMs—0
Domain Adaptation of VLM for Soccer Video Understanding—0
A Challenge to Build Neuro-Symbolic Video AgentsCode0
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?—0
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video UnderstandingCode0
From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations—0
SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation—0
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language ModelsCode0
Gameplay Highlights Generation—0
Seed1.5-VL Technical Report—0
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant—0
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph—0
VideoLLM Benchmarks and Evaluation: A Survey—0
Empowering Agentic Video Analytics Systems with Video Language Models—0
SeriesBench: A Benchmark for Narrative-Driven Drama Series UnderstandingCode0
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation—0
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs—0
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes—0
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection—0
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video UnderstandingCode0
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task—0
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding—0
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?—0
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval—0
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization—0
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild—0
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding—0
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model—0
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking—0
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding—0
How Can Objects Help Video-Language Understanding?—0
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding—0
From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction—0
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models—0
InstructionBench: An Instructional Video Understanding Benchmark—0
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval—0
Moment Quantization for Video Temporal Grounding—0
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding—0
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?—0
Aligned Better, Listen Better for Audio-Visual Large Language Models—0
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description—0
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding—0
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition—0
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts—0
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding—0
Show:102550
← PrevPage 10 of 23Next →

No leaderboard results yet.