SOTAVerified

video narration captioning

Human narration is another critical factor to understand a multi-shot video. It often provides information of the background knowledge and commentator’s view on visual events. We conduct experiments to predict the narration caption of a video-shot and name this task single-shot narration captioning. We adopt the same model structure as single-shot video captioning with the ASR text as additional input, except that the prediction target is the narration caption.

Papers

Showing 11 of 1 papers

TitleStatusHype
Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot VideosCode1
Show:102550

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1OursBLEU-418.8Unverified