SOTAVerified

Video-Text Retrieval

Video-Text retrieval requires understanding of both video and language together. Therefore it's different to video retrieval task.

Papers

Showing 5175 of 111 papers

TitleStatusHype
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalCode0
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectivesCode0
CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive LearningCode0
Diving Deep into the Motion Representation of Video-Text ModelsCode0
Expertized Caption Auto-Enhancement for Video-Text RetrievalCode0
Harvest Video Foundation Models via Efficient Post-PretrainingCode0
Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text RetrievalCode0
Rudder: A Cross Lingual Video and Text Retrieval DatasetCode0
TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible AdapterCode0
Video-Text Retrieval by Supervised Sparse Multi-Grained LearningCode0
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks0
Mask to reconstruct: Cooperative Semantics Completion for Video-text Retrieval0
Masked Contrastive Pre-Training for Efficient Video-Text Retrieval0
LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders0
Leveraging Generative Language Models for Weakly Supervised Sentence Component Analysis in Video-Language Joint Learning0
Rethinking Noisy Video-Text Retrieval via Relation-aware Alignment0
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning0
Retrieving and Highlighting Action with Spatiotemporal Reference0
Learning with Noisy Correspondence0
Learning Context-Adapted Video-Text Retrieval by Attending to User Comments0
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval0
Beyond Coarse-Grained Matching in Video-Text Retrieval0
LaT: Latent Translation with Cycle-Consistency for Video-Text Retrieval0
Stacked Convolutional Deep Encoding Network for Video-Text Retrieval0
HiVLP: Hierarchical Interactive Video-Language Pre-Training0
Show:102550
← PrevPage 3 of 5Next →

No leaderboard results yet.