Contrastive Principal Component Analysis

Abid, Abubakar; Zhang, Martin J.; Bagaria, Vivek K.; Zou, James

Statistics > Machine Learning

arXiv:1709.06716 (stat)

[Submitted on 20 Sep 2017 (v1), last revised 22 Nov 2017 (this version, v2)]

Title:Contrastive Principal Component Analysis

Authors:Abubakar Abid, Martin J. Zhang, Vivek K. Bagaria, James Zou

View PDF

Abstract:We present a new technique called contrastive principal component analysis (cPCA) that is designed to discover low-dimensional structure that is unique to a dataset, or enriched in one dataset relative to other data. The technique is a generalization of standard PCA, for the setting where multiple datasets are available -- e.g. a treatment and a control group, or a mixed versus a homogeneous population -- and the goal is to explore patterns that are specific to one of the datasets. We conduct a wide variety of experiments in which cPCA identifies important dataset-specific patterns that are missed by PCA, demonstrating that it is useful for many applications: subgroup discovery, visualizing trends, feature selection, denoising, and data-dependent standardization. We provide geometrical interpretations of cPCA and show that it satisfies desirable theoretical guarantees. We also extend cPCA to nonlinear settings in the form of kernel cPCA. We have released our code as a python package and documentation is on Github.

Comments:	main body is 10 pages, 9 figures
Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as:	arXiv:1709.06716 [stat.ML]
	(or arXiv:1709.06716v2 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1709.06716

Submission history

From: Abubakar Abid [view email]
[v1] Wed, 20 Sep 2017 03:53:03 UTC (7,201 KB)
[v2] Wed, 22 Nov 2017 00:26:51 UTC (7,201 KB)

Statistics > Machine Learning

Title:Contrastive Principal Component Analysis

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Contrastive Principal Component Analysis

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators