SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 201–250 of 1149 papers

TitleStatusHype
BEARCUBS: A benchmark for computer-using web agents—0
ALLVB: All-in-One Long Video Understanding Benchmark—0
Towards Fine-Grained Video Question Answering—0
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long VideosCode1
Unified Reward Model for Multimodal Understanding and GenerationCode4
EgoLife: Towards Egocentric Life AssistantCode3
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection—0
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation LearningCode1
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models—0
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos—0
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action DetectionCode3
M-LLM Based Video Frame Selection for Efficient Video Understanding—0
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model—0
An Analysis of Data Transformation Effects on Segment Anything 2—0
Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric VideosCode1
Fine-Grained Video Captioning through Scene Graph Consolidation—0
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models—0
AVD2: Accident Video Diffusion for Accident Video Description—0
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval—0
iMOVE: Instance-Motion-Aware Video Understanding—0
VRoPE: Rotary Position Embedding for Video Large Language ModelsCode1
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language ModelCode1
Semantics-aware Test-time Adaptation for 3D Human Pose Estimation—0
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video UnderstandingCode2
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering—0
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis—0
A Survey on Mamba Architecture for Vision Applications—0
CoS: Chain-of-Shot Prompting for Long Video Understanding—0
A Survey on Video Analytics in Cloud-Edge-Terminal Collaborative Systems—0
VideoRoPE: What Makes for Good Video Rotary Position Embedding?Code3
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context AccurayCode3
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs—0
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding—0
A Decade of Action Quality Assessment: Largest Systematic Survey of Trends, Challenges, and Future Directions—0
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesCode1
Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task PerspectivesCode1
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models—0
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context VideosCode7
AIN: The Arabic INclusive Large Multimodal ModelCode2
-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory ConsolidationCode1
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding—0
Understanding Long Videos via LLM-Powered Entity Relation Graphs—0
TinyLLaVA-Video: A Simple Framework of Small-scale Large Multimodal Models for Video UnderstandingCode2
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding—0
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeCode2
Temporal Preference Optimization for Long-Form Video Understanding—0
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video UnderstandingCode5
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling—0
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model—0
MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingCode2
Show:102550
← PrevPage 5 of 23Next →

No leaderboard results yet.