SOTAVerified

Moment Retrieval

Moment retrieval can de defined as the task of "localizing moments in a video given a user query".

Description from: QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries

Image credit: QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries

Papers

Showing 1–25 of 132 papers

TitleStatusHype
DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding—0
Retrieval Augmented Generation Evaluation for Health Documents—0
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection—0
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking—0
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long VideosCode1
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval—0
Moment of Untruth: Dealing with Negative Queries in Video Moment RetrievalCode0
Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection—0
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight DetectionCode1
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models—0
The Devil is in the Spurious Correlation: Boosting Moment Retrieval via Temporal Dynamic Learning—0
A Flexible and Scalable Framework for Video Moment SearchCode1
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight DetectionCode1
DTOS: Dynamic Time Object Sensing with Large Multimodal ModelCode0
Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2DCode0
Length-Aware DETR for Robust Moment RetrievalCode1
DAVE: Diverse Atomic Visual Elements Dataset with High Representation of Vulnerable Road Users in Complex and Unpredictable Environments—0
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning—0
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal GroundingCode1
Agent-based Video Trimming—0
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment RetrievalCode1
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the WildCode1
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment RetrievalCode0
Number it: Temporal Grounding Videos like Flipping MangaCode2
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningCode2
Show:102550
← PrevPage 1 of 6Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1UnLoc-LR@1 IoU=0.566.1—Unverified
2UnLoc-BR@1 IoU=0.564.5—Unverified
3DenoiseLocR@1 IoU=0.559.27—Unverified
4SG-DETR (w/ PT)mAP58.8—Unverified
5SG-DETRmAP54.1—Unverified
6LLaVA-MRmAP52.73—Unverified
7FlashVTGmAP52—Unverified
8InternVideo2-6BmAP49.24—Unverified
9CG-DETR (w/ PT)mAP47.97—Unverified
10VideoLights-B-ptmAP47.94—Unverified
#ModelMetricClaimedVerifiedStatus
1SG-DETR (w/ PT)R@1 IoU=0.571.1—Unverified
2LLaVA-MRR@1 IoU=0.570.65—Unverified
3FlashVTGR@1 IoU=0.570.32—Unverified
4SG-DETRR@1 IoU=0.570.2—Unverified
5InternVideo2-6BR@1 IoU=0.570.03—Unverified
6InternVideo2-1BR@1 IoU=0.568.36—Unverified
7VideoChat-T (FT)R@1 IoU=0.567.1—Unverified
8UniMD+Sync.R@1 IoU=0.563.98—Unverified
9LD-DETRR@1 IoU=0.562.58—Unverified
10VideoLights-B-ptR@1 IoU=0.561.96—Unverified