SOTAVerified

Temporal Sentence Grounding

Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. For this task, different levels of supervision are used. 1) Weak supervision: video-level action category set; 2) Semi-weak supervision: video-level action category set, and action annotations at several timestamps; 3) Full supervision: Action category and action interval annotations of all actions in untrimmed videos.

Papers

Showing 1–10 of 43 papers

TitleStatusHype
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosCode1
Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining—0
Contrast-Unity for Partially-Supervised Temporal Sentence Grounding—0
Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding—0
Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network—0
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language ModelsCode2
Transformer with Controlled Attention for Synchronous Motion CaptioningCode0
Diversifying Query: Region-Guided Transformer for Temporal Sentence GroundingCode0
Video sentence grounding with temporally global textual knowledge—0
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in VideoCode0
Show:102550
← PrevPage 1 of 5Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1DeCafNet-100%R@1,IoU=0.323.2—Unverified
2DeCafNet-50%R@1,IoU=0.321.29—Unverified
3VSLNetR@1,IoU=0.311.7—Unverified