SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 651700 of 1149 papers

TitleStatusHype
DOAD: Decoupled One Stage Action Detection Network0
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering0
Domain Adaptation of VLM for Soccer Video Understanding0
DPMix: Mixture of Depth and Point Cloud Video Experts for 4D Action Segmentation0
Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning0
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model0
DrVideo: Document Retrieval Based Long Video Understanding0
Dilated Temporal Relational Adversarial Network for Generic Video Summarization0
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM0
DualX-VSR: Dual Axial SpatialTemporal Transformer for Real-World Video Super-Resolution without Motion Compensation0
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs0
Dynamic Appearance: A Video Representation for Action Recognition with Joint Training0
Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition0
Dynamic Multistep Reasoning based on Video Scene Graph for Video Question Answering0
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding0
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding0
EAGLE: Egocentric AGgregated Language-video Engine0
Efficient Annotation and Learning for 3D Hand Pose Estimation: A Survey0
Efficient Modelling Across Time of Human Actions and Interactions0
Efficient Motion-Aware Video MLLM0
Efficient Video Understanding via Layered Multi Frame-Rate Analysis0
EgoEnv: Human-centric environment representations from egocentric video0
Egocentric Video Task Translation0
EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding0
Egok360: A 360 Egocentric Kinetic Human Activity Video Dataset0
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation0
ElasticPlay: Interactive Video Summarization with Dynamic Time Budgets0
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding0
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments0
Empowering Agentic Video Analytics Systems with Video Language Models0
End-to-end Generative Pretraining for Multimodal Video Captioning0
End-to-End Joint Semantic Segmentation of Actors and Actions in Video0
End-to-End Video Classification with Knowledge Graphs0
Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning0
Enhancing Long Video Understanding via Hierarchical Event-Based Memory0
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization0
Enhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training0
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis0
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model0
EVA: An Embodied World Model for Future Video Anticipation0
Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment0
EVQAScore: Efficient Video Question Answering Data Evaluation0
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding0
Egocentric and Exocentric Methods: A Short Survey0
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling0
Exploiting Spatial-Temporal Modelling and Multi-Modal Fusion for Human Action Recognition0
Exploring Anchor-based Detection for Ego4D Natural Language Query0
Exploring Missing Modality in Multimodal Egocentric Datasets0
Exploring State Change Capture of Heterogeneous Backbones @ Ego4D Hands and Objects Challenge 20220
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding0
Show:102550
← PrevPage 14 of 23Next →

No leaderboard results yet.