Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 51–100 of 114 papers

Title	Date	Tasks	Status	Hype
Can I Trust Your Answer? Visually Grounded Video Question Answering	Sep 4, 2023	Grounded Video Question AnsweringQuestion Answering	CodeCode Available	1
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection	Aug 29, 2023	DenoisingHighlight Detection	—Unverified	0
Knowing Where to Focus: Event-aware Transformer for Video Grounding	Aug 14, 2023	Moment QueriesSentence	CodeCode Available	1
ViGT: Proposal-free Video Grounding with Learnable Token in Transformer	Aug 11, 2023	Feature Correlationregression	—Unverified	0
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory	Jul 26, 2023	Contrastive LearningVideo Grounding	—Unverified	0
No-frills Temporal Video Grounding: Multi-Scale Neighboring Attention and Zoom-in Boundary Detection	Jul 20, 2023	Boundary DetectionVideo Grounding	—Unverified	0
Dense Video Object Captioning from Disjoint Supervision	Jun 20, 2023	ObjectSentence	—Unverified	0
Boundary-Denoising for Video Activity Localization	Apr 6, 2023	Action DetectionDecoder	CodeCode Available	0
Query-Dependent Video Representation for Moment Retrieval and Highlight Detection	Mar 24, 2023	Highlight DetectionMoment Retrieval	CodeCode Available	2
Generation-Guided Multi-Level Unified Network for Video Grounding	Mar 14, 2023	Video Grounding	—Unverified	0
Text-Visual Prompting for Efficient 2D Temporal Video Grounding	Mar 9, 2023	SentenceVideo Grounding	CodeCode Available	1
Localizing Moments in Long Video Via Multimodal Guidance	Feb 26, 2023	Natural Language Moment RetrievalNatural Language Visual Grounding	CodeCode Available	1
MINOTAUR: Multi-task Video Grounding From Multimodal Queries	Feb 16, 2023	Action DetectionSentence	CodeCode Available	0
Exploiting Auxiliary Caption for Video Grounding	Jan 15, 2023	Contrastive LearningDense Video Captioning	—Unverified	0
WINNER: Weakly-Supervised hIerarchical decompositioN and aligNment for Spatio-tEmporal Video gRounding	Jan 1, 2023	Contrastive LearningSpatio-Temporal Video Grounding	—Unverified	0
Iterative Proposal Refinement for Weakly-Supervised Video Grounding	Jan 1, 2023	SentenceVideo Grounding	—Unverified	0
Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding	Jan 1, 2023	ObjectSpatio-Temporal Video Grounding	—Unverified	0
Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language	Jan 1, 2023	Question AnsweringSelf-Supervised Learning	CodeCode Available	0
Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding	Jan 1, 2023	DecoderSentence	—Unverified	0
A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge	Nov 16, 2022	Action LocalizationNatural Language Queries	CodeCode Available	0
Language-free Training for Zero-shot Video Grounding	Oct 24, 2022	Video Grounding	—Unverified	0
Weakly-Supervised Temporal Article Grounding	Oct 22, 2022	AllArticles	CodeCode Available	1
Graph2Vid: Flow graph to Video Grounding for Weakly-supervised Multi-Step Localization	Oct 10, 2022	Video Grounding	—Unverified	0
On the Effects of Video Grounding on Language Models	Oct 1, 2022	Image CaptioningQuestion Answering	—Unverified	0
Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding	Sep 27, 2022	DecoderSpatio-Temporal Video Grounding	CodeCode Available	1
Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding	Sep 26, 2022	BenchmarkingNatural Language Queries	CodeCode Available	0
CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding	Sep 22, 2022	Contrastive LearningVideo Grounding	CodeCode Available	1
Video-Guided Curriculum Learning for Spoken Video Grounding	Sep 1, 2022	Video Grounding	CodeCode Available	0
Exploiting Feature Diversity for Make-up Temporal Video Grounding	Aug 12, 2022	DiversityVideo Grounding	—Unverified	0
Team PKU-WICT-MIPL PIC Makeup Temporal Video Grounding Challenge 2022 Technical Report	Jul 6, 2022	SentenceTemporal Localization	—Unverified	0
STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding	Jul 6, 2022	Spatio-Temporal Video GroundingVideo Grounding	—Unverified	0
Gaussian Kernel-based Cross Modal Network for Spatio-Temporal Video Grounding	Jul 2, 2022	Spatio-Temporal Video GroundingVideo Grounding	—Unverified	0
Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding	Apr 18, 2022	Action RecognitionAnimal Action Recognition	CodeCode Available	1
Position-aware Location Regression Network for Temporal Video Grounding	Apr 12, 2022	Positionregression	—Unverified	0
TubeDETR: Spatio-Temporal Video Grounding with Transformers	Mar 30, 2022	DecoderLanguage-Based Temporal Localization	CodeCode Available	1
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection	Mar 23, 2022	DecoderHighlight Detection	CodeCode Available	2
End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding	Mar 15, 2022	DescriptiveRepresentation Learning	—Unverified	0
Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding	Mar 8, 2022	Contrastive LearningSentence	—Unverified	0
Explore-And-Match: Bridging Proposal-Based and Proposal-Free With Transformer for Sentence Grounding in Videos	Jan 25, 2022	Natural Language QueriesSentence	CodeCode Available	1
Unsupervised Temporal Video Grounding with Deep Semantic Clustering	Jan 14, 2022	ClusteringSentence	—Unverified	0
Semi-Supervised Video Paragraph Grounding With Contrastive Encoder	Jan 1, 2022	SentenceVideo Grounding	—Unverified	0
Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation	Jan 1, 2022	ObjectReferring Expression Segmentation	—Unverified	0
LocFormer: Enabling Transformers to Perform Temporal Moment Localization on Long Untrimmed Videos With a Feature Sampling Approach	Dec 19, 2021	Inductive BiasVideo Grounding	—Unverified	0
Detecting Moments and Highlights in Videos via Natural Language Queries	Dec 1, 2021	DecoderMoment Retrieval	CodeCode Available	1
End-to-End Dense Video Grounding via Parallel Regression	Sep 23, 2021	regressionSentence	—Unverified	0
On Pursuit of Designing Multi-modal Transformer for Video Grounding	Sep 13, 2021	AllDecoder	—Unverified	0
Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding	Sep 10, 2021	Metric LearningRepresentation Learning	CodeCode Available	1
EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation	Sep 10, 2021	TranslationVideo Grounding	—Unverified	0
Support-Set Based Cross-Supervision for Video Grounding	Aug 24, 2021	Contrastive LearningVideo Grounding	—Unverified	0
VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer	Jul 6, 2021	Image RetrievalKnowledge Distillation	CodeCode Available	1

Show:10 25 50

← PrevPage 2 of 3Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified