SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 401–450 of 1149 papers

TitleStatusHype
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding—0
Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI—0
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments—0
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation—0
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models—0
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges—0
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs—0
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence AlignmentCode0
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs—0
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes—0
Task-Aware KV Compression For Cost-Effective Long Video UnderstandingCode0
PEVLM: Parallel Encoding for Vision-Language Models—0
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning—0
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding—0
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric OptimizationCode0
MambaMia: A State-Space-Model-Based Compression for Efficient Video Understanding in Large Multimodal Models—0
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video UnderstandingCode0
HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person ScenariosCode0
VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks—0
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding—0
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding—0
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding—0
SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis—0
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric VisionCode0
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models—0
DualX-VSR: Dual Axial SpatialTemporal Transformer for Real-World Video Super-Resolution without Motion Compensation—0
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs—0
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval—0
TextVidBench: A Benchmark for Long Video Scene Text Understanding—0
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding—0
METok: Multi-Stage Event-based Token Compression for Efficient Long Video UnderstandingCode0
InterRVOS: Interaction-aware Referring Video Object Segmentation—0
EgoVLM: Policy Optimization for Egocentric Video UnderstandingCode0
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding—0
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding—0
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis—0
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders—0
Learning reusable concepts across different egocentric video understanding tasks—0
VUDG: A Dataset for Video Understanding Domain Generalization—0
Time Blindness: Why Video-Language Models Can't See What Humans Can?—0
ScaleLong: A Multi-Timescale Benchmark for Long Video UnderstandingCode0
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection—0
Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding—0
Universal Visuo-Tactile Video Understanding for Embodied Interaction—0
Two Causally Related Needles in a Video Haystack—0
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic VideosCode0
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models—0
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs—0
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding—0
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game UnderstandingCode0
Show:102550
← PrevPage 9 of 23Next →

No leaderboard results yet.