Accès ouvert

Clustering using shared reference points algorithm based on a sound data model

Article scientifique 2023 Anglais

Résumé

A novel clustering algorithm CSHARP is presented for the purpose of finding clusters of arbitrary shapes and arbitrary densities in high-dimensional feature spaces. It can be considered as a variation of the Shared Nearest Neighbor algorithm (SNN), in which each sample data point votes for the points in its k-nearest neighborhood. Sets of points sharing a common mutual nearest neighbor are considered as dense regions/ blocks. These blocks are the seeds from which clusters may grow. Therefore, CSharp is not a point-to-point clustering algorithm. Rather, it is a block-to-block clustering technique. Much of its advantages come from these facts: Noise points and outliers correspond to blocks of small sizes, and homogeneous blocks highly overlap. The proposed technique is less likely to merge clusters of different densities or different homogeneity. The algorithm has been applied to a variety of low and high-dimensional data sets with superior results over existingtechniques such as DBScan, K-means, Chameleon, Mitosis, and Spectral Clustering. The quality of its results as well as its time complexity, rank it at the front of these techniques.

Citer ce document

Abbas, M., Shoukry, A. (2023). Clustering using shared reference points algorithm based on a sound data model. https://doi.org/10.31219/osf.io/xqv27

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Statistiques

Consultations : 1

Téléchargements : 0