SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 10011050 of 1149 papers

TitleStatusHype
O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning0
OBJECT DYNAMICS DISTILLATION FOR SCENE DECOMPOSITION AND REPRESENTATION0
Occluded Video Instance Segmentation: Dataset and ICCV 2021 Challenge0
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding0
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts0
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks0
OmniTrack: Real-time detection and tracking of objects, text and logos in video0
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding0
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding0
Only Time Can Tell: Discovering Temporal Data for Temporal Modeling0
On the Limitations of Vision-Language Models in Understanding Image Transforms0
Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting0
Open Vocabulary Multi-Label Video Classification0
Open-Vocabulary Spatio-Temporal Action Detection0
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering0
Overview of Tencent Multi-modal Ads Video Understanding Challenge0
Overview of the MedVidQA 2022 Shared Task on Medical Video Question-Answering0
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track0
OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models0
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning0
PcmNet: Position-Sensitive Context Modeling Network for Temporal Action Localization0
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries0
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark0
Personalized Video Summarization by Multimodal Video Understanding0
Person Count Localization in Videos From Noisy Foreground and Detections0
PEVLM: Parallel Encoding for Vision-Language Models0
PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval0
Principles of Visual Tokens for Efficient Video Understanding0
ProBio: A Protocol-guided Multimodal Dataset for Molecular Biology Lab0
Progress-Aware Video Frame Captioning0
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering0
PromptonomyViT: Multi-Task Prompt Learning Improves Video Transformers using Synthetic Scene Data0
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval0
PVChat: Personalized Video Chat with One-Shot Learning0
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models0
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild0
PYSKL: a toolbox for skeleton-based video understanding0
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs0
Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs0
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs0
Query-aware Long Video Localization and Relation Discrimination for Deep Video Understanding0
Query-Conditioned Three-Player Adversarial Network for Video Summarization0
Question Answering is a Format; When is it Useful?0
R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding0
R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding0
Random Temporal Skipping for Multirate Video Analysis0
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph0
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding0
Real-Time Video Highlights for Yahoo Esports0
Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation0
Show:102550
← PrevPage 21 of 23Next →

No leaderboard results yet.