AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant

2022-03-08Code Available1· sign in to hype

Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, Mike Zheng Shou

Code Available — Be the first to reproduce this paper.

Code

github.com/showlab/Q2A
Officialpytorch★ 23
github.com/starsholic/loveu-cvpr22-aqtc
pytorch★ 6
github.com/jaykim9870/cvpr-22_loveu_unipyler
pytorch★ 2
github.com/zcfinal/loveu-cvpr23-aqtc
pytorch★ 1

Abstract

A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave for 1 minute?". However, there is still no clear task definition and suitable benchmarks. In this paper, we define a new task called Affordance-centric Question-driven Task Completion, where the AI assistant should learn from instructional videos to provide step-by-step help in the user's view. To support the task, we constructed AssistQ, a new dataset comprising 531 question-answer samples from 100 newly filmed instructional videos. We also developed a novel Question-to-Actions (Q2A) model to address the AQTC task and validate it on the AssistQ dataset. The results show that our model significantly outperforms several VQA-related baselines while still having large room for improvement. We expect our task and dataset to advance Egocentric AI Assistant's development. Our project page is available at: https://showlab.github.io/assistq/.

Tasks

Visual Question Answering (VQA)

AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant

Code

Abstract

Tasks

Reproductions