SOTAVerified

Video Understanding

A crucial task of Video Understanding is to recognise and localise (in space and time) different actions or events appearing in the video.

Source: Action Detection from a Robot-Car Perspective

Papers

Showing 10111020 of 1149 papers

TitleStatusHype
On the Limitations of Vision-Language Models in Understanding Image Transforms0
Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting0
Open Vocabulary Multi-Label Video Classification0
Open-Vocabulary Spatio-Temporal Action Detection0
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering0
Overview of Tencent Multi-modal Ads Video Understanding Challenge0
Overview of the MedVidQA 2022 Shared Task on Medical Video Question-Answering0
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track0
OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models0
OW-VISCapTor: Abstractors for Open-World Video Instance Segmentation and Captioning0
Show:102550
← PrevPage 102 of 115Next →

No leaderboard results yet.