A Complex Networks Approach for Data Clustering

Rodrigues, Francisco A.; de Arruda, Guilherme Ferraz; Costa, Luciano da Fontoura

Physics > Data Analysis, Statistics and Probability

arXiv:1101.5141 (physics)

[Submitted on 26 Jan 2011]

Title:A Complex Networks Approach for Data Clustering

Authors:Francisco A. Rodrigues, Guilherme Ferraz de Arruda, Luciano da Fontoura Costa

View PDF

Abstract:Many methods have been developed for data clustering, such as k-means, expectation maximization and algorithms based on graph theory. In this latter case, graphs are generally constructed by taking into account the Euclidian distance as a similarity measure, and partitioned using spectral methods. However, these methods are not accurate when the clusters are not well separated. In addition, it is not possible to automatically determine the number of clusters. These limitations can be overcome by taking into account network community identification algorithms. In this work, we propose a methodology for data clustering based on complex networks theory. We compare different metrics for quantifying the similarity between objects and take into account three community finding techniques. This approach is applied to two real-world databases and to two sets of artificially generated data. By comparing our method with traditional clustering approaches, we verify that the proximity measures given by the Chebyshev and Manhattan distances are the most suitable metrics to quantify the similarity between objects. In addition, the community identification method based on the greedy optimization provides the smallest misclassification rates.

Comments:	9 pages, 8 Figures
Subjects:	Data Analysis, Statistics and Probability (physics.data-an); Machine Learning (cs.LG); Social and Information Networks (cs.SI); Physics and Society (physics.soc-ph)
Cite as:	arXiv:1101.5141 [physics.data-an]
	(or arXiv:1101.5141v1 [physics.data-an] for this version)
	https://doi.org/10.48550/arXiv.1101.5141

Submission history

From: Francisco Aparecido Rodrigues [view email]
[v1] Wed, 26 Jan 2011 19:58:58 UTC (3,128 KB)

Physics > Data Analysis, Statistics and Probability

Title:A Complex Networks Approach for Data Clustering

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Physics > Data Analysis, Statistics and Probability

Title:A Complex Networks Approach for Data Clustering

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators