Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

Wu, Qi; Wang, Peng; Shen, Chunhua; Reid, Ian; Hengel, Anton van den

Computer Science > Computer Vision and Pattern Recognition

arXiv:1711.07613 (cs)

[Submitted on 21 Nov 2017]

Title:Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

Authors:Qi Wu, Peng Wang, Chunhua Shen, Ian Reid, Anton van den Hengel

View PDF

Abstract:The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialogue that has taken place. The key challenge in Visual Dialogue is thus maintaining a consistent, and natural dialogue while continuing to answer questions correctly. We present a novel approach that combines Reinforcement Learning and Generative Adversarial Networks (GANs) to generate more human-like responses to questions. The GAN helps overcome the relative paucity of training data, and the tendency of the typical MLE-based approach to generate overly terse answers. Critically, the GAN is tightly integrated into the attention mechanism that generates human-interpretable reasons for each answer. This means that the discriminative model of the GAN has the task of assessing whether a candidate answer is generated by a human or not, given the provided reason. This is significant because it drives the generative model to produce high quality answers that are well supported by the associated reasoning. The method also generates the state-of-the-art results on the primary benchmark.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:1711.07613 [cs.CV]
	(or arXiv:1711.07613v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1711.07613

Submission history

From: Qi Wu [view email]
[v1] Tue, 21 Nov 2017 03:11:49 UTC (3,169 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CV

< prev | next >

new | recent | 2017-11

Change to browse by:

cs
cs.AI
cs.CL

References & Citations

DBLP - CS Bibliography

listing | bibtex

Qi Wu
Peng Wang
Chunhua Shen
Ian D. Reid
Anton van den Hengel

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators