Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation

Bhojanapalli, Srinadh; Chakrabarti, Ayan; Jain, Himanshu; Kumar, Sanjiv; Lukasik, Michal; Veit, Andreas

Computer Science > Machine Learning

arXiv:2106.08823 (cs)

[Submitted on 16 Jun 2021]

Title:Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation

Authors:Srinadh Bhojanapalli, Ayan Chakrabarti, Himanshu Jain, Sanjiv Kumar, Michal Lukasik, Andreas Veit

View PDF

Abstract:State-of-the-art transformer models use pairwise dot-product based self-attention, which comes at a computational cost quadratic in the input sequence length. In this paper, we investigate the global structure of attention scores computed using this dot product mechanism on a typical distribution of inputs, and study the principal components of their variation. Through eigen analysis of full attention score matrices, as well as of their individual rows, we find that most of the variation among attention scores lie in a low-dimensional eigenspace. Moreover, we find significant overlap between these eigenspaces for different layers and even different transformer models. Based on this, we propose to compute scores only for a partial subset of token pairs, and use them to estimate scores for the remaining pairs. Beyond investigating the accuracy of reconstructing attention scores themselves, we investigate training transformer models that employ these approximations, and analyze the effect on overall accuracy. Our analysis and the proposed method provide insights into how to balance the benefits of exact pair-wise attention and its significant computational expense.

Comments:	14 pages
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2106.08823 [cs.LG]
	(or arXiv:2106.08823v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2106.08823

Submission history

From: Srinadh Bhojanapalli [view email]
[v1] Wed, 16 Jun 2021 14:38:42 UTC (3,887 KB)

Computer Science > Machine Learning

Title:Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators