SOTAVerified

Moment Retrieval

Moment retrieval can de defined as the task of "localizing moments in a video given a user query".

Description from: QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries

Image credit: QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries

Papers

Showing 2650 of 132 papers

TitleStatusHype
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video UnderstandingCode1
Saliency-Guided DETR for Moment Retrieval and Highlight DetectionCode1
Show and Guide: Instructional-Plan Grounded Vision and Language ModelCode0
EAGLE: Egocentric AGgregated Language-video Engine0
Language-based Audio Moment RetrievalCode3
D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching0
QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval0
Disentangle and denoise: Tackling context misalignment for video moment retrieval0
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight DetectionCode3
SLVideo: A Sign Language Video Moment Retrieval Framework0
Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment RetrievalCode2
Multi-sentence Video Grounding for Long Video Generation0
EA-VTR: Event-Aware Video-Text Retrieval0
TVR-Ranking: A Dataset for Ranked Video Moment Retrieval with Imprecise QueriesCode0
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal GroundingCode0
The Surprising Effectiveness of Multimodal Large Language Models for Video Moment RetrievalCode2
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval0
2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval0
Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels0
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingCode2
Context-Enhanced Video Moment Retrieval with Large Language Models0
MLP: Motion Label Prior for Temporal Sentence Localization in Untrimmed 3D Human MotionsCode1
Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight DetectionCode1
UniMD: Towards Unifying Moment Retrieval and Temporal Action DetectionCode2
R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal GroundingCode0
Show:102550
← PrevPage 2 of 6Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1UnLoc-LR@1 IoU=0.566.1Unverified
2UnLoc-BR@1 IoU=0.564.5Unverified
3DenoiseLocR@1 IoU=0.559.27Unverified
4SG-DETR (w/ PT)mAP58.8Unverified
5SG-DETRmAP54.1Unverified
6LLaVA-MRmAP52.73Unverified
7FlashVTGmAP52Unverified
8InternVideo2-6BmAP49.24Unverified
9CG-DETR (w/ PT)mAP47.97Unverified
10VideoLights-B-ptmAP47.94Unverified
#ModelMetricClaimedVerifiedStatus
1SG-DETR (w/ PT)R@1 IoU=0.571.1Unverified
2LLaVA-MRR@1 IoU=0.570.65Unverified
3FlashVTGR@1 IoU=0.570.32Unverified
4SG-DETRR@1 IoU=0.570.2Unverified
5InternVideo2-6BR@1 IoU=0.570.03Unverified
6InternVideo2-1BR@1 IoU=0.568.36Unverified
7VideoChat-T (FT)R@1 IoU=0.567.1Unverified
8UniMD+Sync.R@1 IoU=0.563.98Unverified
9LD-DETRR@1 IoU=0.562.58Unverified
10VideoLights-B-ptR@1 IoU=0.561.96Unverified