SOTAVerified

Moment Retrieval

Moment retrieval can de defined as the task of "localizing moments in a video given a user query".

Description from: QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries

Image credit: QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries

Papers

Showing 51–100 of 132 papers

TitleStatusHype
Deconfounded Video Moment Retrieval with Causal InterventionCode1
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight DetectionCode1
Background-aware Moment Detection for Video Moment RetrievalCode1
Partially Relevant Video RetrievalCode1
Video Moment Retrieval from Text Queries via Single Frame AnnotationCode1
Retrieval Augmented Generation Evaluation for Health Documents—0
2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval—0
Agent-based Video Trimming—0
A Survey on Video Moment Localization—0
AxIoU: An Axiomatically Justified Measure for Video Moment Retrieval—0
Coarse to Fine: Video Retrieval before Moment Localization—0
Context-Enhanced Video Moment Retrieval with Large Language Models—0
Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval—0
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments—0
DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding—0
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection—0
Disentangle and denoise: Tackling context misalignment for video moment retrieval—0
D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching—0
EAGLE: Egocentric AGgregated Language-video Engine—0
EA-VTR: Event-Aware Video-Text Retrieval—0
Event-aware Video Corpus Moment Retrieval—0
Faster Video Moment Retrieval with Point-Level Supervision—0
Fast Video Moment Retrieval—0
FedVMR: A New Federated Learning method for Video Moment Retrieval—0
Generating Adjacency Matrix for Video Relocalization—0
Generative Video Diffusion for Unseen Cross-Domain Video Moment Retrieval—0
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features—0
Graph Neural Network for Video Relocalization—0
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection—0
Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels—0
Interactive Video Corpus Moment Retrieval using Reinforcement Learning—0
Language Guided Networks for Cross-modal Moment Retrieval—0
Leveraging Generative Language Models for Weakly Supervised Sentence Component Analysis in Video-Language Joint Learning—0
MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment—0
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval—0
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval—0
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval—0
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection—0
Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval—0
Multi-scale 2D Representation Learning for weakly-supervised moment retrieval—0
Multi-sentence Video Grounding for Long Video Generation—0
Multi-video Moment Ranking with Multimodal Clue—0
QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval—0
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning—0
R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding—0
R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding—0
SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment Retrieval—0
SLVideo: A Sign Language Video Moment Retrieval Framework—0
Temporal Perceiving Video-Language Pre-training—0
Text-based Localization of Moments in a Video Corpus—0
Show:102550
← PrevPage 2 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1UnLoc-LR@1 IoU=0.566.1—Unverified
2UnLoc-BR@1 IoU=0.564.5—Unverified
3DenoiseLocR@1 IoU=0.559.27—Unverified
4SG-DETR (w/ PT)mAP58.8—Unverified
5SG-DETRmAP54.1—Unverified
6LLaVA-MRmAP52.73—Unverified
7FlashVTGmAP52—Unverified
8InternVideo2-6BmAP49.24—Unverified
9CG-DETR (w/ PT)mAP47.97—Unverified
10VideoLights-B-ptmAP47.94—Unverified
#ModelMetricClaimedVerifiedStatus
1SG-DETR (w/ PT)R@1 IoU=0.571.1—Unverified
2LLaVA-MRR@1 IoU=0.570.65—Unverified
3FlashVTGR@1 IoU=0.570.32—Unverified
4SG-DETRR@1 IoU=0.570.2—Unverified
5InternVideo2-6BR@1 IoU=0.570.03—Unverified
6InternVideo2-1BR@1 IoU=0.568.36—Unverified
7VideoChat-T (FT)R@1 IoU=0.567.1—Unverified
8UniMD+Sync.R@1 IoU=0.563.98—Unverified
9LD-DETRR@1 IoU=0.562.58—Unverified
10VideoLights-B-ptR@1 IoU=0.561.96—Unverified