SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 681690 of 1149 papers

TitleStatusHype
End-to-end Generative Pretraining for Multimodal Video Captioning0
End-to-End Joint Semantic Segmentation of Actors and Actions in Video0
End-to-End Video Classification with Knowledge Graphs0
Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning0
Enhancing Long Video Understanding via Hierarchical Event-Based Memory0
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization0
Enhancing Transformer for Video Understanding Using Gated Multi-Level Attention and Temporal Adversarial Training0
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis0
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model0
EVA: An Embodied World Model for Future Video Anticipation0
Show:102550
← PrevPage 69 of 115Next →

No leaderboard results yet.