SOTAVerified

Video Retrieval

The objective of video retrieval is as follows: given a text query and a pool of candidate videos, select the video which corresponds to the text query. Typically, the videos are returned as a ranked list of candidates and scored via document retrieval metrics.

Papers

Showing 301–350 of 486 papers

TitleStatusHype
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling—0
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval—0
Multi-Granularity Graph Pooling for Video-based Person Re-Identification—0
Multimodal Approach for Video Surveillance Indexing and Retrieval—0
Multimodal Contextualized Support for Enhancing Video Retrieval System—0
Multimodal Skip-gram Using Convolutional Pseudowords—0
Multiple Visual-Semantic Embedding for Video Retrieval from Query Sentence—0
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval—0
MultiVENT: Multilingual Videos of Events with Aligned Natural Text—0
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions—0
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality—0
Near-duplicate video detection featuring coupled temporal and perceptual visual structures and logical inference based matching—0
Neighborhood Preserving Hashing for Scalable Video Retrieval—0
Neural Graph Matching for Video Retrieval in Large-Scale Video-driven E-commerce—0
NEWSKVQA: Knowledge-Aware News Video Question Answering—0
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision—0
Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval—0
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks—0
Perfect Match in Video Retrieval—0
PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval—0
PolySmart @ TRECVid 2024 Medical Video Question Answering—0
Pose-Aided Video-based Person Re-Identification via Recurrent Graph Convolutional Network—0
Probabilistic Representations for Video Contrastive Learning—0
ProTA: Probabilistic Token Aggregation for Text-Video Retrieval—0
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval—0
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval—0
QSAM-Net: Rain streak removal by quaternion neural network with self-attention module—0
Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model—0
Query by Semantic Sketch—0
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning—0
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter—0
Real-time analysis of cataract surgery videos using statistical models—0
Renmin University of China at TRECVID 2022: Improving Video Search by Feature Fusion and Negation Understanding—0
RNNs, CNNs and Transformers in Human Action Recognition: A Survey and a Hybrid Model—0
Self-supervised Spatiotemporal Representation Learning by Exploiting Video Continuity—0
Self-supervised Temporal Learning—0
Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder—0
Self-Supervised Video Representation Learning with Meta-Contrastive Network—0
Self-Supervised Video Representation Learning by Video Incoherence Detection—0
Self-supervised Video Retrieval Transformer Network—0
Semantic Image Retrieval by Uniting Deep Neural Networks and Cognitive Architectures—0
Semantic Video Entity Linking Based on Visual Content and Metadata—0
Semantic Video Moments Retrieval at Scale: A New Task and a Baseline—0
Semi-automatic Data Annotation System for Multi-Target Multi-Camera Vehicle Tracking—0
Sharing Hash Codes for Multiple Purposes—0
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval—0
Sign Language Video Retrieval with Free-Form Textual Queries—0
Sinkhorn Transformations for Single-Query Postprocessing in Text-Video Retrieval—0
SketchGAN: Joint Sketch Completion and Recognition With Generative Adversarial Network—0
SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training—0
Show:102550
← PrevPage 7 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1OmniVectext-to-video R@1089.4—Unverified
2CLIP4Cliptext-to-video R@1081.6—Unverified
3OmniVec (pretrained)text-to-video R@1078.6—Unverified
4HunYuan_tvr (huge)text-to-video R@162.9—Unverified
5CLIP-ViPtext-to-video R@157.7—Unverified
6PIDRotext-to-video R@155.9—Unverified
7DMAE (ViT-B/16)text-to-video R@155.5—Unverified
8HunYuan_tvrtext-to-video R@155—Unverified
9MuLTItext-to-video R@154.7—Unverified
10EERCFtext-to-video R@154.1—Unverified
#ModelMetricClaimedVerifiedStatus
1Aurora (ours, r=64)text-to-video R@577.4—Unverified
2InternVideo2-6Btext-to-video R@174.2—Unverified
3vid-TLDR (UMT-L)text-to-video R@172.3—Unverified
4VASTtext-to-video R@172—Unverified
5COSAtext-to-video R@170.5—Unverified
6UMT-L (ViT-L/16)text-to-video R@170.4—Unverified
7GRAMtext-to-video R@167.3—Unverified
8VALORtext-to-video R@161.5—Unverified
9TESTA (ViT-B/16)text-to-video R@161.2—Unverified
10VindLUtext-to-video R@161.2—Unverified
#ModelMetricClaimedVerifiedStatus
1GRAMtext-to-video R@164—Unverified
2VASTtext-to-video R@163.9—Unverified
3InternVideo2-6Btext-to-video R@162.8—Unverified
4VALORtext-to-video R@159.9—Unverified
5UMT-L (ViT-L/16)text-to-video R@158.8—Unverified
6vid-TLDR (UMT-L)text-to-video R@158.1—Unverified
7COSAtext-to-video R@157.9—Unverified
8InternVideo2-6Btext-to-video R@155.9—Unverified
9InternVideotext-to-video R@155.2—Unverified
10VLABtext-to-video R@155.1—Unverified
#ModelMetricClaimedVerifiedStatus
1EMCL-Net (Ours)++ LSMDC Rohrbach et al. (2015)text-to-video R@1053.7—Unverified
2InternVideo2-6Btext-to-video R@146.4—Unverified
3vid-TLDR (UMT-L)text-to-video R@143.1—Unverified
4UMT-L (ViT-L/16)text-to-video R@143—Unverified
5HunYuan_tvr (huge)text-to-video R@140.4—Unverified
6COSAtext-to-video R@139.4—Unverified
7mPLUG-2text-to-video R@134.4—Unverified
8VALORtext-to-video R@134.2—Unverified
9InternVideotext-to-video R@134—Unverified
10InternVideo2-6Btext-to-video R@133.8—Unverified