Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 26–50 of 114 papers

Title	Date	Tasks	Status	Hype
Text-Visual Prompting for Efficient 2D Temporal Video Grounding	Mar 9, 2023	SentenceVideo Grounding	CodeCode Available	1
Localizing Moments in Long Video Via Multimodal Guidance	Feb 26, 2023	Natural Language Moment RetrievalNatural Language Visual Grounding	CodeCode Available	1
Weakly-Supervised Temporal Article Grounding	Oct 22, 2022	AllArticles	CodeCode Available	1
Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding	Sep 27, 2022	DecoderSpatio-Temporal Video Grounding	CodeCode Available	1
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding	Sep 22, 2022	Contrastive LearningVideo Grounding	CodeCode Available	1
Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding	Apr 18, 2022	Action RecognitionAnimal Action Recognition	CodeCode Available	1
TubeDETR: Spatio-Temporal Video Grounding with Transformers	Mar 30, 2022	DecoderLanguage-Based Temporal Localization	CodeCode Available	1
Explore-And-Match: Bridging Proposal-Based and Proposal-Free With Transformer for Sentence Grounding in Videos	Jan 25, 2022	Natural Language QueriesSentence	CodeCode Available	1
Detecting Moments and Highlights in Videos via Natural Language Queries	Dec 1, 2021	DecoderMoment Retrieval	CodeCode Available	1
Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding	Sep 10, 2021	Metric LearningRepresentation Learning	CodeCode Available	1
VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer	Jul 6, 2021	Image RetrievalKnowledge Distillation	CodeCode Available	1
VLG-Net: Video-Language Graph Matching Network for Video Grounding	Nov 19, 2020	Graph MatchingMoment Retrieval	CodeCode Available	1
Human-centric Spatio-Temporal Video Grounding With Visual Transformers	Nov 10, 2020	Referring ExpressionSentence	CodeCode Available	1
Dense Regression Network for Video Grounding	Apr 7, 2020	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1
Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences	Jan 19, 2020	FormObject	CodeCode Available	1
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding	Jul 17, 2025	Video GroundingVideo Understanding	—Unverified	0
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models	May 24, 2025	BenchmarkingVideo Grounding	—Unverified	0
Enhancing Weakly Supervised Video Grounding via Diverse Inference Strategies for Boundary and Prediction Selection	Mar 29, 2025	PredictionVideo Grounding	—Unverified	0
VideoGEM: Training-free Action Grounding in Videos	Mar 26, 2025	Video Grounding	—Unverified	0
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability	Mar 18, 2025	Language ModelingLanguage Modelling	—Unverified	0
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding	Jan 28, 2025	object-detectionObject Detection	—Unverified	0
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding	Jan 1, 2025	Action UnderstandingSpatio-Temporal Video Grounding	—Unverified	0
Consistency of Compositional Generalization across Multiple Levels	Dec 18, 2024	Meta-LearningQuestion Answering	CodeCode Available	0
Multi-Scale Contrastive Learning for Video Temporal Grounding	Dec 10, 2024	Contrastive LearningData Augmentation	—Unverified	0
Video LLMs for Temporal Reasoning in Long Videos	Dec 4, 2024	Action SegmentationDense Video Captioning	—Unverified	0

Show:10 25 50

← PrevPage 2 of 5Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified