SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 9511000 of 1149 papers

TitleStatusHype
M^33D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding0
M^3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action Recognition0
MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection0
Making Every Frame Matter: Continuous Video Understanding for Large Models via Adaptive State Modeling0
MAMBA4D: Efficient Long-Sequence Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models0
MambaMia: A State-Space-Model-Based Compression for Efficient Video Understanding in Large Multimodal Models0
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations0
Massively Parallel Video Networks0
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model0
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding0
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal Localization0
Memory Consolidation Enables Long-Context Video Understanding0
Memory-enhanced Retrieval Augmentation for Long Video Understanding0
Memory-Guided Semantic Learning Network for Temporal Sentence Grounding0
MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD0
MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound0
Mid-level Representation for Visual Recognition0
Mimic The Raw Domain: Accelerating Action Recognition in the Compressed Domain0
M-LLM Based Video Frame Selection for Efficient Video Understanding0
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding0
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning0
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding0
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning0
MM-Ego: Towards Building Egocentric Multimodal LLMs0
Moment Quantization for Video Temporal Grounding0
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval0
Morph: Flexible Acceleration for 3D CNN-based Video Understanding0
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models0
Motion-Guided Masking for Spatiotemporal Representation Learning0
Motion Sensitive Contrastive Learning for Self-supervised Video Representation0
MovieLLM: Enhancing Long Video Understanding with AI-Generated Movies0
MovieNet: A Holistic Dataset for Movie Understanding0
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning0
MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding0
MRSN: Multi-Relation Support Network for Video Action Detection0
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language0
Multi-kernel learning of deep convolutional features for action recognition0
Multimodal High-order Relation Transformer for Scene Boundary Detection0
Multimodal Intent Discovery from Livestream Videos0
Multi-modal Representation Learning for Video Advertisement Content Structuring0
Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation0
Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding0
Multi-scale 2D Temporal Map Diffusion Models for Natural Language Video Localization0
Multi-Scale Contrastive Learning for Video Temporal Grounding0
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding0
Multiview Transformers for Video Recognition0
MVTamperBench: Evaluating Robustness of Vision-Language Models0
Representation Learning on Visual-Symbolic Graphs for Video Understanding0
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision0
Non-local NetVLAD Encoding for Video Classification0
Show:102550
← PrevPage 20 of 23Next →

No leaderboard results yet.