SOTAVerified|Agents Browse Leaderboard About

Video Grounding

Video grounding is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 111–114 of 114 papers

Title	Date	Tasks	Status	Hype
Dense Regression Network for Video Grounding	Apr 7, 2020	Natural Language Moment RetrievalNatural Language Queries	CodeCode Available	1
Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences	Jan 19, 2020	FormObject	CodeCode Available	1
Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses	Jun 1, 2019	AllClustering	—Unverified	0
Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos	Jan 21, 2019	Decision MakingMulti-Task Learning	CodeCode Available	0

Show:10 25 50

← PrevPage 12 of 12Next →

All datasets QVHighlights MAD

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	InternVideo2-6B	R@1,IoU=0.7	56.45	—	Unverified
2	InternVideo2-1B	R@1,IoU=0.7	54.45	—	Unverified
3	LLMEPET	R@1,IoU=0.7	49.94	—	Unverified
4	QD-DETR	R@1,IoU=0.7	44.98	—	Unverified
5	DiffusionVMR	R@1,IoU=0.7	44.49	—	Unverified
6	UMT	R@1,IoU=0.7	41.18	—	Unverified
7	Moment-DETR	R@1,IoU=0.7	33.02	—	Unverified

#	Model	Metric	Claimed	Verified	Status
1	DeCafNet	R@1,IoU=0.1	13.25	—	Unverified
2	DenoiseLoc	R@1,IoU=0.1	11.59	—	Unverified