Generating Video Description using Sequence-to-sequence Model with Temporal Attention

2016-12-01COLING 2016Unverified0· sign in to hype

Natsuda Laokulrat, Sang Phan, Noriki Nishida, Raphael Shu, Yo Ehara, Naoaki Okazaki, Yusuke Miyao, Hideki Nakayama

Unverified — Be the first to reproduce this paper.

Abstract

Automatic video description generation has recently been getting attention after rapid advancement in image caption generation. Automatically generating description for a video is more challenging than for an image due to its temporal dynamics of frames. Most of the work relied on Recurrent Neural Network (RNN) and recently attentional mechanisms have also been applied to make the model learn to focus on some frames of the video while generating each word in a describing sentence. In this paper, we focus on a sequence-to-sequence approach with temporal attention mechanism. We analyze and compare the results from different attention model configuration. By applying the temporal attention mechanism to the system, we can achieve a METEOR score of 0.310 on Microsoft Video Description dataset, which outperformed the state-of-the-art system so far.

Tasks

Caption Generation Sentence Video Classification Video Description

Generating Video Description using Sequence-to-sequence Model with Temporal Attention

Abstract

Tasks

Reproductions