Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 11–20 of 114 papers

Title	Date	Tasks	Status	Hype	Score
Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval	Jul 21, 2024	General KnowledgeHighlight Detection	CodeCode Available	2	5
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency	Jun 2, 2025	reinforcement-learningReinforcement Learning	CodeCode Available	2	5
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding	Feb 16, 2025	AttributeObject	CodeCode Available	1	5
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding	Sep 22, 2022	Contrastive LearningVideo Grounding	CodeCode Available	1	5
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding	Dec 27, 2023	SentenceTemporal Sentence Grounding	CodeCode Available	1	5
Knowing Where to Focus: Event-aware Transformer for Video Grounding	Aug 14, 2023	Moment QueriesSentence	CodeCode Available	1	5
Localizing Moments in Long Video Via Multimodal Guidance	Feb 26, 2023	Natural Language Moment RetrievalNatural Language Visual Grounding	CodeCode Available	1	5
HawkEye: Training Video-Text LLMs for Grounding Text in Videos	Mar 15, 2024	Video GroundingVideo Question Answering	CodeCode Available	1	5
Dense Regression Network for Video Grounding	Apr 7, 2020	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1	5
Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding	Sep 27, 2022	DecoderSpatio-Temporal Video Grounding	CodeCode Available	1	5

Show:10 25 50

← PrevPage 2 of 12Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified