SOTAVerified

Text to Video Retrieval

She's gone I can't find her anywhere I'm looking everywhere for her Everywhere is dark

Papers

Showing 51–75 of 75 papers

TitleStatusHype
Efficient End-to-End Video Question Answering with Pyramidal Multimodal TransformerCode0
Fighting FIRe with FIRE: Assessing the Validity of Text-to-Video Retrieval Benchmarks—0
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval—0
Retrieving and Highlighting Action with Spatiotemporal Reference—0
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners—0
Learning text-to-video retrieval from image captioning—0
Learning Trajectory-Word Alignments for Video-Language Tasks—0
Sakuga-42M Dataset: Scaling Up Cartoon Research—0
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review—0
Leveraging Generative Language Models for Weakly Supervised Sentence Component Analysis in Video-Language Joint Learning—0
MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization—0
SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training—0
Support-set bottlenecks for video-text representation learning—0
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval—0
Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment—0
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval—0
COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval—0
CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning—0
TeachCLIP: Multi-Grained Teaching for Efficient Text-to-Video Retrieval—0
Distilling Vision-Language Models on Millions of Videos—0
Temporal Perceiving Video-Language Pre-training—0
EA-VTR: Event-Aware Video-Text Retrieval—0
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval—0
An Empirical Study of Frame Selection for Text-to-Video Retrieval—0
E-ViLM: Efficient Video-Language Model via Masked Video Modeling with Semantic Vector-Quantized Tokenizer—0
Show:102550
← PrevPage 3 of 3Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1FROZEN-revisedmAP23.39—Unverified
2FROZEN-revised (two-stream)text-to-video R@112.8—Unverified
#ModelMetricClaimedVerifiedStatus
1CLIP4Cliptext-to-video R@144.5—Unverified
#ModelMetricClaimedVerifiedStatus
1X-CLIP (Cross-Lingual)R@132.3—Unverified