Understanding Attention for Vision-and-Language Tasks

Cao, Feiqi; Han, Soyeon Caren; Long, Siqu; Xu, Changwei; Poon, Josiah

Computer Science > Computer Vision and Pattern Recognition

arXiv:2208.08104 (cs)

[Submitted on 17 Aug 2022 (v1), last revised 22 Sep 2022 (this version, v2)]

Title:Understanding Attention for Vision-and-Language Tasks

Authors:Feiqi Cao, Soyeon Caren Han, Siqu Long, Changwei Xu, Josiah Poon

View PDF

Abstract:Attention mechanism has been used as an important component across Vision-and-Language(VL) tasks in order to bridge the semantic gap between visual and textual features. While attention has been widely used in VL tasks, it has not been examined the capability of different attention alignment calculation in bridging the semantic gap between visual and textual clues. In this research, we conduct a comprehensive analysis on understanding the role of attention alignment by looking into the attention score calculation methods and check how it actually represents the visual region's and textual token's significance for the global assessment. We also analyse the conditions which attention score calculation mechanism would be more (or less) interpretable, and which may impact the model performance on three different VL tasks, including visual question answering, text-to-image generation, text-and-image matching (both sentence and image retrieval). Our analysis is the first of its kind and provides useful insights of the importance of each attention alignment score calculation when applied at the training phase of VL tasks, commonly ignored in attention-based cross modal models, and/or pretrained models. Our code is available at: this https URL

Comments:	Accepted in COLING 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as:	arXiv:2208.08104 [cs.CV]
	(or arXiv:2208.08104v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2208.08104

Submission history

From: Feiqi Cao [view email]
[v1] Wed, 17 Aug 2022 06:45:07 UTC (4,248 KB)
[v2] Thu, 22 Sep 2022 06:24:44 UTC (4,249 KB)

✅2024-10-01: arxiv.org is back to normal.✅

Computer Science > Computer Vision and Pattern Recognition

Title:Understanding Attention for Vision-and-Language Tasks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

✅2024-10-01: arxiv.org is back to normal.✅

Computer Science > Computer Vision and Pattern Recognition

Title:Understanding Attention for Vision-and-Language Tasks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators