SOTAVerified

Video Summarization

Video Summarization aims to generate a short synopsis that summarizes the video content by selecting its most informative and important parts. The produced summary is usually composed of a set of representative video frames (a.k.a. video key-frames), or video fragments (a.k.a. video key-fragments) that have been stitched in chronological order to form a shorter video. The former type of a video summary is known as video storyboard, and the latter type is known as video skim.

Source: Video Summarization Using Deep Neural Networks: A Survey Image credit: iJRASET

Papers

Showing 101–125 of 280 papers

TitleStatusHype
A Memory Network Approach for Story-Based Temporal Summarization of 360° Videos—0
CSTA: CNN-based Spatiotemporal Attention for Video Summarization—0
HSA-RNN: Hierarchical Structure-Adaptive RNN for Video Summarization—0
Creating Summaries from User Videos—0
How Local is the Local Diversity? Reinforcing Sequential Determinantal Point Processes with Dynamic Ground Sets for Supervised Video Summarization—0
How Good is a Video Summary? A New Benchmarking Dataset and Evaluation Framework Towards Realistic Video Summarization—0
Highlight Detection With Pairwise Deep Ranking for First-Person Video Summarization—0
Human Pose Estimation using Motion Priors and Ensemble Models—0
Co-Regularized Deep Representations for Video Summarization—0
Image Conditioned Keyframe-Based Video Summarization Using Object Detection—0
A Paradigm for Building Generalized Models of Human Image Perception Through Data Fusion—0
Hierarchical Recurrent Neural Network for Video Summarization—0
Hierarchical Multimodal Transformer to Summarize Videos—0
Group Activity Recognition by Using Effective Multiple Modality Relation Representation With Temporal-Spatial Attention—0
Conditional Modeling Based Automatic Video Summarization—0
A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts—0
Global-and-Local Relative Position Embedding for Unsupervised Video Summarization—0
Key Frame Extraction with Attention Based Deep Neural Networks—0
Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video—0
Large-Margin Determinantal Point Processes—0
Large Model based Sequential Keyframe Extraction for Video Summarization—0
Large-Scale Video Summarization Using Web-Image Priors—0
Generating Natural Language Summaries for Multimedia—0
Comprehensive Video Understanding: Video summarization with content-based video recommender design—0
Gaze-Enabled Egocentric Video Summarization via Constrained Submodular Maximization—0
Show:102550
← PrevPage 5 of 12Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1PGL-SUMF1-score (Canonical)55.6—Unverified
2RR-STGF1-score (Canonical)54.5—Unverified
3DSNetF1-score (Canonical)53—Unverified
4VASNetF1-score (Canonical)49.71—Unverified
5M-AVSF1-score (Canonical)44.4—Unverified
6CSTAKendall's Tau0.25—Unverified
#ModelMetricClaimedVerifiedStatus
1RR-STGF1-score (Canonical)63—Unverified
2DSNetF1-score (Canonical)62.1—Unverified
3VASNetF1-score (Canonical)61.42—Unverified
4PGL-SUMF1-score (Canonical)61—Unverified
5M-AVSF1-score (Canonical)61—Unverified
6CSTAKendall's Tau0.19—Unverified
#ModelMetricClaimedVerifiedStatus
1Shotluck-Holmes (3.1B)CIDEr152.3—Unverified
2Shotluck-Holmes (3.1B)CIDEr63.2—Unverified
3SUM-shotCIDEr8.6—Unverified
#ModelMetricClaimedVerifiedStatus
1EgoVLPv2F1 (avg)52.08—Unverified
2EgoVLPF1 (avg)49.72—Unverified
#ModelMetricClaimedVerifiedStatus
1PGL-SUMMAP (50%)61.6—Unverified
#ModelMetricClaimedVerifiedStatus
1VTSUM-BLIP1 shot Micro-F123.5—Unverified