Fusion Subspace Clustering: Full and Incomplete Data

Pimentel-Alarcón, Daniel L.; Mahmood, Usman

Computer Science > Machine Learning

arXiv:1808.00628 (cs)

[Submitted on 2 Aug 2018]

Title:Fusion Subspace Clustering: Full and Incomplete Data

Authors:Daniel L. Pimentel-Alarcón, Usman Mahmood

View PDF

Abstract:Modern inference and learning often hinge on identifying low-dimensional structures that approximate large scale data. Subspace clustering achieves this through a union of linear subspaces. However, in contemporary applications data is increasingly often incomplete, rendering standard (full-data) methods inapplicable. On the other hand, existing incomplete-data methods present major drawbacks, like lifting an already high-dimensional problem, or requiring a super polynomial number of samples. Motivated by this, we introduce a new subspace clustering algorithm inspired by fusion penalties. The main idea is to permanently assign each datum to a subspace of its own, and minimize the distance between the subspaces of all data, so that subspaces of the same cluster get fused together. Our approach is entirely new to both, full and missing data, and unlike other methods, it directly allows noise, it requires no liftings, it allows low, high, and even full-rank data, it approaches optimal (information-theoretic) sampling rates, and it does not rely on other methods such as low-rank matrix completion to handle missing data. Furthermore, our extensive experiments on both real and synthetic data show that our approach performs comparably to the state-of-the-art with complete data, and dramatically better if data is missing.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1808.00628 [cs.LG]
	(or arXiv:1808.00628v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1808.00628

Submission history

From: Daniel L. Pimentel-Alarcón [view email]
[v1] Thu, 2 Aug 2018 01:54:15 UTC (405 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2018-08

Change to browse by:

cs
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Daniel L. Pimentel-Alarcón
Usman Mahmood

export BibTeX citation

Computer Science > Machine Learning

Title:Fusion Subspace Clustering: Full and Incomplete Data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Fusion Subspace Clustering: Full and Incomplete Data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators