Aligned Dual Channel Graph Convolutional Network for Visual Question Answering

2020-07-01ACL 2020Unverified0· sign in to hype

Qingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng, Junying Chen, Ho-fung Leung, Qing Li

Unverified — Be the first to reproduce this paper.

Abstract

Visual question answering aims to answer the natural language question about a given image. Existing graph-based methods only focus on the relations between objects in an image and neglect the importance of the syntactic dependency relations between words in a question. To simultaneously capture the relations between objects in an image and the syntactic dependency relations between words in a question, we propose a novel dual channel graph convolutional network (DC-GCN) for better combining visual and textual advantages. The DC-GCN model consists of three parts: an I-GCN module to capture the relations between objects in an image, a Q-GCN module to capture the syntactic dependency relations between words in a question, and an attention alignment module to align image representations and question representations. Experimental results show that our model achieves comparable performance with the state-of-the-art approaches.

Tasks

Question Answering Visual Question Answering Visual Question Answering (VQA)

Aligned Dual Channel Graph Convolutional Network for Visual Question Answering

Abstract

Tasks

Reproductions