SOTAVerified

Spatio-Temporal Video Grounding

Spatio-temporal video grounding is a computer vision and natural language processing (NLP) task that involves linking textual descriptions to specific spatio-temporal regions or moments in a video. In other words, it aims to determine which parts of a video correspond to a given textual query or description. This task is essential for various applications, including video summarization, content-based video retrieval, video captioning, and more.

Papers

Showing 21–22 of 22 papers

TitleStatusHype
STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding—0
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding—0
Show:102550
← PrevPage 3 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1TA-STVGVal m_vIoU40.2—Unverified
2CG-STVGVal m_vIoU39.5—Unverified
3STVGFormerVal m_vIoU38.7—Unverified
4TubeDETRVal m_vIoU36.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TA-STVGm_vIoU39.1—Unverified
2CG-STVGm_vIoU38.4—Unverified
3TubeDETRm_vIoU32.4—Unverified
#ModelMetricClaimedVerifiedStatus
1TA-STVGDeclarative m_vIoU34.4—Unverified
2CG-STVGDeclarative m_vIoU34—Unverified
3TubeDETRDeclarative m_vIoU30.4—Unverified