Learning to Answer Questions From Image Using Convolutional Neural   Network

Lin Ma; Zhengdong Lu; Hang Li

arXiv:1506.00333·cs.CL·November 16, 2015·29 cites

Learning to Answer Questions From Image Using Convolutional Neural Network

Lin Ma, Zhengdong Lu, Hang Li

PDF

Open Access

TL;DR

This paper introduces a convolutional neural network framework for image question answering, effectively encoding images and questions and learning their interactions to improve answer accuracy.

Contribution

The paper presents a novel end-to-end CNN model that jointly encodes images and questions for improved image QA performance.

Findings

01

Significantly outperforms previous state-of-the-art on DAQUAR dataset.

02

Achieves superior results on COCO-QA benchmark.

03

Demonstrates the effectiveness of multimodal CNN architecture.

Abstract

In this paper, we propose to employ the convolutional neural network (CNN) for the image question answering (QA). Our proposed CNN provides an end-to-end framework with convolutional architectures for learning not only the image and question representations, but also their inter-modal interactions to produce the answer. More specifically, our model consists of three CNNs: one image CNN to encode the image content, one sentence CNN to compose the words of the question, and one multimodal convolution layer to learn their joint representation for the classification in the space of candidate answer words. We demonstrate the efficacy of our proposed model on the DAQUAR and COCO-QA datasets, which are two benchmark datasets for the image QA, with the performances significantly outperforming the state-of-the-art.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Domain Adaptation and Few-Shot Learning · Topic Modeling

MethodsConvolution