Open-Ended Multi-Modal Relational Reasoning for Video Question Answering

2020-12-01Code Available0· sign in to hype

Haozheng Luo, Ruiyang Qin, Chenwei Xu, Guo Ye, Zening Luo

Code Available — Be the first to reproduce this paper.

Code

github.com/robinzixuan/Video-Question-Answering-HRI
Officialpytorch★ 1

Abstract

In this paper, we introduce a robotic agent specifically designed to analyze external environments and address participants' questions. The primary focus of this agent is to assist individuals using language-based interactions within video-based scenes. Our proposed method integrates video recognition technology and natural language processing models within the robotic agent. We investigate the crucial factors affecting human-robot interactions by examining pertinent issues arising between participants and robot agents. Methodologically, our experimental findings reveal a positive relationship between trust and interaction efficiency. Furthermore, our model demonstrates a 2\% to 3\% performance enhancement in comparison to other benchmark methods.

Tasks

Question Answering Relational Reasoning Video Question Answering Video Recognition Visual Question Answering (VQA)

Open-Ended Multi-Modal Relational Reasoning for Video Question Answering

Code

Abstract

Tasks

Reproductions