SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 125 of 1149 papers

TitleStatusHype
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding0
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New BenchmarksCode1
Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI0
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments0
Omni-Video: Democratizing Unified Video Understanding and GenerationCode2
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models0
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video UnderstandingCode1
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation0
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges0
Kwai Keye-VL Technical ReportCode4
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs0
GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement LearningCode7
Flash-VStream: Efficient Real-Time Understanding for Long Video StreamsCode3
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence AlignmentCode0
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs0
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMsCode2
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes0
Task-Aware KV Compression For Cost-Effective Long Video UnderstandingCode0
PEVLM: Parallel Encoding for Vision-Language Models0
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning0
video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language ModelsCode2
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding0
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric OptimizationCode0
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video UnderstandingCode0
MambaMia: A State-Space-Model-Based Compression for Efficient Video Understanding in Large Multimodal Models0
Show:102550
← PrevPage 1 of 46Next →

No leaderboard results yet.