Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 26–50 of 114 papers

Title	Date	Tasks	Status	Hype
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding	Sep 22, 2022	Contrastive LearningVideo Grounding	CodeCode Available	1
Dense Regression Network for Video Grounding	Apr 7, 2020	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos	Mar 9, 2025	Action LocalizationBoundary Detection	CodeCode Available	1
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format	Nov 27, 2024	Dense Video CaptioningGrounded Video Question Answering	CodeCode Available	1
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding	Dec 27, 2023	SentenceTemporal Sentence Grounding	CodeCode Available	1
Can I Trust Your Answer? Visually Grounded Video Question Answering	Sep 4, 2023	Grounded Video Question AnsweringQuestion Answering	CodeCode Available	1
Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding	Sep 10, 2021	Metric LearningRepresentation Learning	CodeCode Available	1
Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding	Apr 18, 2022	Action RecognitionAnimal Action Recognition	CodeCode Available	1
Object-Shot Enhanced Grounding Network for Egocentric Video	May 7, 2025	Video Grounding	CodeCode Available	1
Explore-And-Match: Bridging Proposal-Based and Proposal-Free With Transformer for Sentence Grounding in Videos	Jan 25, 2022	Natural Language QueriesSentence	CodeCode Available	1
Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection	Nov 28, 2023	Contrastive LearningHighlight Detection	CodeCode Available	1
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding	Mar 13, 2025	ObjectVideo Grounding	CodeCode Available	1
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos	May 22, 2025	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1
VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer	Jul 6, 2021	Image RetrievalKnowledge Distillation	CodeCode Available	1
Localizing Moments in Long Video Via Multimodal Guidance	Feb 26, 2023	Natural Language Moment RetrievalNatural Language Visual Grounding	CodeCode Available	1
EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation	Sep 10, 2021	TranslationVideo Grounding	—Unverified	0
EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model	Dec 5, 2023	Boundary DetectionLanguage Modeling	—Unverified	0
Enhancing Weakly Supervised Video Grounding via Diverse Inference Strategies for Boundary and Prediction Selection	Mar 29, 2025	PredictionVideo Grounding	—Unverified	0
End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding	Mar 15, 2022	DescriptiveRepresentation Learning	—Unverified	0
Artemis: Towards Referential Understanding in Complex Videos	Jun 1, 2024	Text SummarizationVideo Grounding	—Unverified	0
Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding	Dec 21, 2023	Domain AdaptationUnsupervised Domain Adaptation	—Unverified	0
End-to-End Dense Video Grounding via Parallel Regression	Sep 23, 2021	regressionSentence	—Unverified	0
Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding	Jan 1, 2023	ObjectSpatio-Temporal Video Grounding	—Unverified	0
Iterative Proposal Refinement for Weakly-Supervised Video Grounding	Jan 1, 2023	SentenceVideo Grounding	—Unverified	0
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection	Aug 29, 2023	DenoisingHighlight Detection	—Unverified	0

Show:10 25 50

← PrevPage 2 of 5Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified