UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog

Chen, Cheng; Zhu, Yudong; Tan, Zhenshan; Cheng, Qingrong; Jiang, Xin; Liu, Qun; Gu, Xiaodong

Computer Science > Computer Vision and Pattern Recognition

arXiv:2205.00423 (cs)

[Submitted on 1 May 2022 (v1), last revised 3 May 2022 (this version, v2)]

Title:UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog

Authors:Cheng Chen, Yudong Zhu, Zhenshan Tan, Qingrong Cheng, Xin Jiang, Qun Liu, Xiaodong Gu

View PDF

Abstract:Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two complementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1.

Comments:	Accepted in CVPR 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2205.00423 [cs.CV]
	(or arXiv:2205.00423v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2205.00423

Submission history

From: Chen Cheng [view email]
[v1] Sun, 1 May 2022 08:36:18 UTC (15,178 KB)
[v2] Tue, 3 May 2022 09:25:00 UTC (4,807 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators