Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 1–50 of 114 papers

Title	Date	Tasks	Status	Hype	Score
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding	Mar 22, 2024	Action ClassificationAction Recognition	CodeCode Available	7	5
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding	Jan 14, 2025	Embodied Question AnsweringHallucination	CodeCode Available	4	5
SnAG: Scalable and Accurate Video Grounding	Apr 2, 2024	Video GroundingVideo Understanding	CodeCode Available	4	5
PG-Video-LLaVA: Pixel Grounding Large Video-Language Models	Nov 22, 2023	BenchmarkingPhrase Grounding	CodeCode Available	2	5
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency	Jun 2, 2025	reinforcement-learningReinforcement Learning	CodeCode Available	2	5
Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval	Jul 21, 2024	General KnowledgeHighlight Detection	CodeCode Available	2	5
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection	Mar 23, 2022	DecoderHighlight Detection	CodeCode Available	2	5
VTimeLLM: Empower LLM to Grasp Video Moments	Nov 30, 2023	Dense Video CaptioningTemporal Relation Extraction	CodeCode Available	2	5
Query-Dependent Video Representation for Moment Retrieval and Highlight Detection	Mar 24, 2023	Highlight DetectionMoment Retrieval	CodeCode Available	2	5
TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM	Mar 17, 2025	Video Grounding	CodeCode Available	2	5
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding	Jan 14, 2025	Feature CompressionLanguage Modeling	CodeCode Available	2	5
Context-Guided Spatio-Temporal Video Grounding	Jan 3, 2024	ObjectSpatio-Temporal Video Grounding	CodeCode Available	2	5
TubeDETR: Spatio-Temporal Video Grounding with Transformers	Mar 30, 2022	DecoderLanguage-Based Temporal Localization	CodeCode Available	1	5
HawkEye: Training Video-Text LLMs for Grounding Text in Videos	Mar 15, 2024	Video GroundingVideo Question Answering	CodeCode Available	1	5
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos	Mar 9, 2025	Action LocalizationBoundary Detection	CodeCode Available	1	5
Text-Visual Prompting for Efficient 2D Temporal Video Grounding	Mar 9, 2023	SentenceVideo Grounding	CodeCode Available	1	5
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning	Jan 12, 2025	Dense Video CaptioningVideo Captioning	CodeCode Available	1	5
Detecting Moments and Highlights in Videos via Natural Language Queries	Dec 1, 2021	DecoderMoment Retrieval	CodeCode Available	1	5
Human-centric Spatio-Temporal Video Grounding With Visual Transformers	Nov 10, 2020	Referring ExpressionSentence	CodeCode Available	1	5
Grounded Question-Answering in Long Egocentric Videos	Dec 11, 2023	Video GroundingVideo Question Answering	CodeCode Available	1	5
Dense Regression Network for Video Grounding	Apr 7, 2020	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1	5
Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences	Jan 19, 2020	FormObject	CodeCode Available	1	5
Weakly-Supervised Temporal Article Grounding	Oct 22, 2022	AllArticles	CodeCode Available	1	5
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding	Feb 16, 2025	AttributeObject	CodeCode Available	1	5
Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding	Sep 27, 2022	DecoderSpatio-Temporal Video Grounding	CodeCode Available	1	5
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding	Sep 22, 2022	Contrastive LearningVideo Grounding	CodeCode Available	1	5
Knowing Where to Focus: Event-aware Transformer for Video Grounding	Aug 14, 2023	Moment QueriesSentence	CodeCode Available	1	5
VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer	Jul 6, 2021	Image RetrievalKnowledge Distillation	CodeCode Available	1	5
Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding	Sep 10, 2021	Metric LearningRepresentation Learning	CodeCode Available	1	5
VLG-Net: Video-Language Graph Matching Network for Video Grounding	Nov 19, 2020	Graph MatchingMoment Retrieval	CodeCode Available	1	5
Object-Shot Enhanced Grounding Network for Egocentric Video	May 7, 2025	Video Grounding	CodeCode Available	1	5
Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding	Apr 18, 2022	Action RecognitionAnimal Action Recognition	CodeCode Available	1	5
Explore-And-Match: Bridging Proposal-Based and Proposal-Free With Transformer for Sentence Grounding in Videos	Jan 25, 2022	Natural Language QueriesSentence	CodeCode Available	1	5
Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection	Nov 28, 2023	Contrastive LearningHighlight Detection	CodeCode Available	1	5
Can I Trust Your Answer? Visually Grounded Video Question Answering	Sep 4, 2023	Grounded Video Question AnsweringQuestion Answering	CodeCode Available	1	5
Localizing Moments in Long Video Via Multimodal Guidance	Feb 26, 2023	Natural Language Moment RetrievalNatural Language Visual Grounding	CodeCode Available	1	5
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding	Dec 27, 2023	SentenceTemporal Sentence Grounding	CodeCode Available	1	5
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos	May 22, 2025	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1	5
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format	Nov 27, 2024	Dense Video CaptioningGrounded Video Question Answering	CodeCode Available	1	5
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding	Mar 13, 2025	ObjectVideo Grounding	CodeCode Available	1	5
Boundary-Denoising for Video Activity Localization	Apr 6, 2023	Action DetectionDecoder	CodeCode Available	0	5
Consistency of Compositional Generalization across Multiple Levels	Dec 18, 2024	Meta-LearningQuestion Answering	CodeCode Available	0	5
Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding	Sep 26, 2022	BenchmarkingNatural Language Queries	CodeCode Available	0	5
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding	Mar 21, 2024	Video Grounding	CodeCode Available	0	5
Dual-Path Temporal Map Optimization for Make-up Temporal Video Grounding	Sep 12, 2023	Sentencetext similarity	CodeCode Available	0	5
Artemis: Towards Referential Understanding in Complex Videos	Jun 1, 2024	Text SummarizationVideo Grounding	CodeCode Available	0	5
Interventional Video Grounding with Dual Contrastive Learning	Jun 21, 2021	Causal InferenceContrastive Learning	CodeCode Available	0	5
Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos	Jan 21, 2019	Decision MakingMulti-Task Learning	CodeCode Available	0	5
Dense Video Object Captioning from Disjoint Supervision	Jun 20, 2023	ObjectSentence	CodeCode Available	0	5
A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge	Nov 16, 2022	Action LocalizationNatural Language Queries	CodeCode Available	0	5

Show:10 25 50

← PrevPage 1 of 3Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified