Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 81–90 of 114 papers

Title	Date	Tasks	Status
Exploring the Effect of Primitives for Compositional Generalization in Vision-and-Language	Jan 1, 2023	Question AnsweringSelf-Supervised Learning	CodeCode Available
Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding	Jan 1, 2023	DecoderSentence	—Unverified
Iterative Proposal Refinement for Weakly-Supervised Video Grounding	Jan 1, 2023	SentenceVideo Grounding	—Unverified
A Simple Transformer-Based Model for Ego4D Natural Language Queries Challenge	Nov 16, 2022	Action LocalizationNatural Language Queries	CodeCode Available
Language-free Training for Zero-shot Video Grounding	Oct 24, 2022	Video Grounding	—Unverified
Graph2Vid: Flow graph to Video Grounding for Weakly-supervised Multi-Step Localization	Oct 10, 2022	Video Grounding	—Unverified
On the Effects of Video Grounding on Language Models	Oct 1, 2022	Image CaptioningQuestion Answering	—Unverified
Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding	Sep 26, 2022	BenchmarkingNatural Language Queries	CodeCode Available
Video-Guided Curriculum Learning for Spoken Video Grounding	Sep 1, 2022	Video Grounding	CodeCode Available
Exploiting Feature Diversity for Make-up Temporal Video Grounding	Aug 12, 2022	DiversityVideo Grounding	—Unverified

Show:10 25 50

← PrevPage 9 of 12Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified