A statistical interpretation of spectral embedding: the generalised random dot product graph

Rubin-Delanchy, Patrick; Cape, Joshua; Tang, Minh; Priebe, Carey E.

Statistics > Machine Learning

arXiv:1709.05506 (stat)

[Submitted on 16 Sep 2017 (v1), last revised 16 Nov 2021 (this version, v5)]

Title:A statistical interpretation of spectral embedding: the generalised random dot product graph

Authors:Patrick Rubin-Delanchy, Joshua Cape, Minh Tang, Carey E. Priebe

View PDF

Abstract:Spectral embedding is a procedure which can be used to obtain vector representations of the nodes of a graph. This paper proposes a generalisation of the latent position network model known as the random dot product graph, to allow interpretation of those vector representations as latent position estimates. The generalisation is needed to model heterophilic connectivity (e.g., `opposites attract') and to cope with negative eigenvalues more generally. We show that, whether the adjacency or normalised Laplacian matrix is used, spectral embedding produces uniformly consistent latent position estimates with asymptotically Gaussian error (up to identifiability). The standard and mixed membership stochastic block models are special cases in which the latent positions take only $K$ distinct vector values, representing communities, or live in the $(K-1)$-simplex with those vertices, respectively. Under the stochastic block model, our theory suggests spectral clustering using a Gaussian mixture model (rather than $K$-means) and, under mixed membership, fitting the minimum volume enclosing simplex, existing recommendations previously only supported under non-negative-definite assumptions. Empirical improvements in link prediction (over the random dot product graph), and the potential to uncover richer latent structure (than posited under the standard or mixed membership stochastic block models) are demonstrated in a cyber-security example.

Comments:	34 pages; 12 figures
Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
MSC classes:	62H30, 62H12, 62E20,
Cite as:	arXiv:1709.05506 [stat.ML]
	(or arXiv:1709.05506v5 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1709.05506

Submission history

From: Patrick Rubin-Delanchy Dr [view email]
[v1] Sat, 16 Sep 2017 12:30:40 UTC (198 KB)
[v2] Thu, 21 Sep 2017 12:58:52 UTC (199 KB)
[v3] Sun, 29 Jul 2018 19:02:30 UTC (646 KB)
[v4] Wed, 8 Jan 2020 16:27:46 UTC (1,017 KB)
[v5] Tue, 16 Nov 2021 11:17:52 UTC (3,035 KB)

Statistics > Machine Learning

Title:A statistical interpretation of spectral embedding: the generalised random dot product graph

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:A statistical interpretation of spectral embedding: the generalised random dot product graph

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators