Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 114 papers

Title	Date	Tasks	Status	Score
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding	Mar 21, 2024	Video Grounding	CodeCode Available	5
Cross-Modal learning for Audio-Visual Video Parsing	Apr 3, 2021	Event DetectionMultiple Instance Learning	CodeCode Available	5
Cross-modal Contrastive Learning with Asymmetric Co-attention Network for Video Moment Retrieval	Dec 12, 2023	Contrastive LearningMoment Retrieval	CodeCode Available	5
Consistency of Compositional Generalization across Multiple Levels	Dec 18, 2024	Meta-LearningQuestion Answering	CodeCode Available	5
VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding	Jan 1, 2024	Spatio-Temporal Video GroundingVideo Grounding	—Unverified	0
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding	Jul 17, 2025	Video GroundingVideo Understanding	—Unverified	0
Video LLMs for Temporal Reasoning in Long Videos	Dec 4, 2024	Action SegmentationDense Video Captioning	—Unverified	0
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition	May 7, 2024	Large Language ModelMultimodal Large Language Model	—Unverified	0
ViGT: Proposal-free Video Grounding with Learnable Token in Transformer	Aug 11, 2023	Feature Correlationregression	—Unverified	0
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses	Aug 3, 2024	Natural Language QueriesVideo Grounding	—Unverified	0
WINNER: Weakly-Supervised hIerarchical decompositioN and aligNment for Spatio-tEmporal Video gRounding	Jan 1, 2023	Contrastive LearningSpatio-Temporal Video Grounding	—Unverified	0
Artemis: Towards Referential Understanding in Complex Videos	Jun 1, 2024	Text SummarizationVideo Grounding	—Unverified	0
Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding	Jun 20, 2021	Spatio-Temporal Video GroundingVideo Grounding	—Unverified	0
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding	Jun 11, 2024	regressionVideo Grounding	—Unverified	0
Cascaded Prediction Network via Segment Tree for Temporal Video Grounding	Jun 19, 2021	SentenceVideo Grounding	—Unverified	0
Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos	Mar 23, 2021	Referring ExpressionReferring Expression Comprehension	—Unverified	0
Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding	Jan 1, 2023	ObjectSpatio-Temporal Video Grounding	—Unverified	0
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding	Jan 28, 2025	object-detectionObject Detection	—Unverified	0
Dense Video Object Captioning from Disjoint Supervision	Jun 20, 2023	ObjectSentence	—Unverified	0
Described Spatial-Temporal Video Detection	Jul 8, 2024	Multi-class ClassificationTemporal Localization	—Unverified	0
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection	Aug 29, 2023	DenoisingHighlight Detection	—Unverified	0
End-to-End Dense Video Grounding via Parallel Regression	Sep 23, 2021	regressionSentence	—Unverified	0
End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding	Mar 15, 2022	DescriptiveRepresentation Learning	—Unverified	0
Enhancing Weakly Supervised Video Grounding via Diverse Inference Strategies for Boundary and Prediction Selection	Mar 29, 2025	PredictionVideo Grounding	—Unverified	0
EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model	Dec 5, 2023	Boundary DetectionLanguage Modeling	—Unverified	0
EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation	Sep 10, 2021	TranslationVideo Grounding	—Unverified	0
Exploiting Feature Diversity for Make-up Temporal Video Grounding	Aug 12, 2022	DiversityVideo Grounding	—Unverified	0
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory	Jul 26, 2023	Contrastive LearningVideo Grounding	—Unverified	0
Gaussian Kernel-based Cross Modal Network for Spatio-Temporal Video Grounding	Jul 2, 2022	Spatio-Temporal Video GroundingVideo Grounding	—Unverified	0
Exploiting Auxiliary Caption for Video Grounding	Jan 15, 2023	Contrastive LearningDense Video Captioning	—Unverified	0
Generation-Guided Multi-Level Unified Network for Video Grounding	Mar 14, 2023	Video Grounding	—Unverified	0
Graph2Vid: Flow graph to Video Grounding for Weakly-supervised Multi-Step Localization	Oct 10, 2022	Video Grounding	—Unverified	0
Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding	Jan 1, 2023	DecoderSentence	—Unverified	0
Iterative Proposal Refinement for Weakly-Supervised Video Grounding	Jan 1, 2023	SentenceVideo Grounding	—Unverified	0
Language-free Training for Zero-shot Video Grounding	Oct 24, 2022	Video Grounding	—Unverified	0
LLM4VG: Large Language Models Evaluation for Video Grounding	Dec 21, 2023	Image CaptioningVideo Grounding	—Unverified	0
LocFormer: Enabling Transformers to Perform Temporal Moment Localization on Long Untrimmed Videos With a Feature Sampling Approach	Dec 19, 2021	Inductive BiasVideo Grounding	—Unverified	0
Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation	Jan 1, 2022	ObjectReferring Expression Segmentation	—Unverified	0
Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding	Dec 21, 2023	Domain AdaptationUnsupervised Domain Adaptation	—Unverified	0
Multi-Scale Contrastive Learning for Video Temporal Grounding	Dec 10, 2024	Contrastive LearningData Augmentation	—Unverified	0
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding	Mar 8, 2022	Contrastive LearningSentence	—Unverified	0
Multi-sentence Video Grounding for Long Video Generation	Jul 18, 2024	Moment RetrievalRetrieval	—Unverified	0
No-frills Temporal Video Grounding: Multi-Scale Neighboring Attention and Zoom-in Boundary Detection	Jul 20, 2023	Boundary DetectionVideo Grounding	—Unverified	0
Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses	Jun 1, 2019	AllClustering	—Unverified	0
Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding	Aug 16, 2020	DiversityObject	—Unverified	0
On Pursuit of Designing Multi-modal Transformer for Video Grounding	Sep 13, 2021	AllDecoder	—Unverified	0
On the Effects of Video Grounding on Language Models	Oct 1, 2022	Image CaptioningQuestion Answering	—Unverified	0
Parallel Attention Network with Sequence Matching for Video Grounding	May 18, 2021	Representation LearningVideo Grounding	—Unverified	0
Position-aware Location Regression Network for Temporal Video Grounding	Apr 12, 2022	Positionregression	—Unverified	0
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models	May 24, 2025	BenchmarkingVideo Grounding	—Unverified	0

Show:10 25 50

← PrevPage 2 of 3Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified