Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 114 papers

Title	Date	Tasks	Status
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding	Nov 25, 2024	Dense Video CaptioningTransfer Learning	—Unverified
SimBase: A Simple Baseline for Temporal Video Grounding	Nov 12, 2024	Video Grounding	—Unverified
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses	Aug 3, 2024	Natural Language QueriesVideo Grounding	—Unverified
Multi-sentence Video Grounding for Long Video Generation	Jul 18, 2024	Moment RetrievalRetrieval	—Unverified
Described Spatial-Temporal Video Detection	Jul 8, 2024	Multi-class ClassificationTemporal Localization	—Unverified
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding	Jun 11, 2024	regressionVideo Grounding	—Unverified
Simplify Implant Depth Prediction as Video Grounding: A Texture Perceive Implant Depth Prediction Network	Jun 7, 2024	Depth EstimationDepth Prediction	—Unverified
Artemis: Towards Referential Understanding in Complex Videos	Jun 1, 2024	Text SummarizationVideo Grounding	CodeCode Available
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition	May 7, 2024	Large Language ModelMultimodal Large Language Model	—Unverified
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding	Apr 1, 2024	MambaState Space Models	—Unverified
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding	Mar 21, 2024	Video Grounding	CodeCode Available
VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding	Jan 1, 2024	Spatio-Temporal Video GroundingVideo Grounding	—Unverified
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding	Dec 31, 2023	Spatio-Temporal Video GroundingVideo Grounding	—Unverified
Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding	Dec 21, 2023	Domain AdaptationUnsupervised Domain Adaptation	—Unverified
LLM4VG: Large Language Models Evaluation for Video Grounding	Dec 21, 2023	Image CaptioningVideo Grounding	—Unverified
Cross-modal Contrastive Learning with Asymmetric Co-attention Network for Video Moment Retrieval	Dec 12, 2023	Contrastive LearningMoment Retrieval	CodeCode Available
EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model	Dec 5, 2023	Boundary DetectionLanguage Modeling	—Unverified
Exploring Iterative Refinement with Diffusion Models for Video Grounding	Oct 26, 2023	SentenceVideo Grounding	CodeCode Available
Dual-Path Temporal Map Optimization for Make-up Temporal Video Grounding	Sep 12, 2023	Sentencetext similarity	CodeCode Available
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection	Aug 29, 2023	DenoisingHighlight Detection	—Unverified
ViGT: Proposal-free Video Grounding with Learnable Token in Transformer	Aug 11, 2023	Feature Correlationregression	—Unverified
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory	Jul 26, 2023	Contrastive LearningVideo Grounding	—Unverified
No-frills Temporal Video Grounding: Multi-Scale Neighboring Attention and Zoom-in Boundary Detection	Jul 20, 2023	Boundary DetectionVideo Grounding	—Unverified
Dense Video Object Captioning from Disjoint Supervision	Jun 20, 2023	ObjectSentence	CodeCode Available
Boundary-Denoising for Video Activity Localization	Apr 6, 2023	Action DetectionDecoder	CodeCode Available
Generation-Guided Multi-Level Unified Network for Video Grounding	Mar 14, 2023	Video Grounding	—Unverified
MINOTAUR: Multi-task Video Grounding From Multimodal Queries	Feb 16, 2023	Action DetectionSentence	CodeCode Available
Exploiting Auxiliary Caption for Video Grounding	Jan 15, 2023	Contrastive LearningDense Video Captioning	—Unverified
WINNER: Weakly-Supervised hIerarchical decompositioN and aligNment for Spatio-tEmporal Video gRounding	Jan 1, 2023	Contrastive LearningSpatio-Temporal Video Grounding	—Unverified
Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding	Jan 1, 2023	ObjectSpatio-Temporal Video Grounding	—Unverified
Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language	Jan 1, 2023	Question AnsweringSelf-Supervised Learning	CodeCode Available
Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding	Jan 1, 2023	DecoderSentence	—Unverified
Iterative Proposal Refinement for Weakly-Supervised Video Grounding	Jan 1, 2023	SentenceVideo Grounding	—Unverified
A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge	Nov 16, 2022	Action LocalizationNatural Language Queries	CodeCode Available
Language-free Training for Zero-shot Video Grounding	Oct 24, 2022	Video Grounding	—Unverified
Graph2Vid: Flow graph to Video Grounding for Weakly-supervised Multi-Step Localization	Oct 10, 2022	Video Grounding	—Unverified
On the Effects of Video Grounding on Language Models	Oct 1, 2022	Image CaptioningQuestion Answering	—Unverified
Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding	Sep 26, 2022	BenchmarkingNatural Language Queries	CodeCode Available
Video-Guided Curriculum Learning for Spoken Video Grounding	Sep 1, 2022	Video Grounding	CodeCode Available
Exploiting Feature Diversity for Make-up Temporal Video Grounding	Aug 12, 2022	DiversityVideo Grounding	—Unverified
Team PKU-WICT-MIPL PIC Makeup Temporal Video Grounding Challenge 2022 Technical Report	Jul 6, 2022	SentenceTemporal Localization	—Unverified
STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding	Jul 6, 2022	Spatio-Temporal Video GroundingVideo Grounding	—Unverified
Gaussian Kernel-based Cross Modal Network for Spatio-Temporal Video Grounding	Jul 2, 2022	Spatio-Temporal Video GroundingVideo Grounding	—Unverified
Position-aware Location Regression Network for Temporal Video Grounding	Apr 12, 2022	Positionregression	—Unverified
End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding	Mar 15, 2022	DescriptiveRepresentation Learning	—Unverified
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding	Mar 8, 2022	Contrastive LearningSentence	—Unverified
Unsupervised Temporal Video Grounding with Deep Semantic Clustering	Jan 14, 2022	ClusteringSentence	—Unverified
Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation	Jan 1, 2022	ObjectReferring Expression Segmentation	—Unverified
Semi-Supervised Video Paragraph Grounding With Contrastive Encoder	Jan 1, 2022	SentenceVideo Grounding	—Unverified
LocFormer: Enabling Transformers to Perform Temporal Moment Localization on Long Untrimmed Videos With a Feature Sampling Approach	Dec 19, 2021	Inductive BiasVideo Grounding	—Unverified

Show:10 25 50

← PrevPage 2 of 3Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified