How many clusters exist? Answer via maximum clustering similarity implemented in R

Finding the number of clusters in a data set is considered as one of the fundamental problems in cluster analysis. This paper integrates maximum clustering similarity (MCS), for finding the optimal number of clusters, into R©statistical software through the package MCSim. The similarity between the...

وصف كامل

محفوظ في:
التفاصيل البيبلوغرافية
المؤلف الرئيسي: Zogheib, Bashar (author)
منشور في: 2019
الوصول للمادة أونلاين:https://doi.org/10.1080/24709360.2019.1615770
https://dspace.auk.edu.kw/handle/11675/5752
الوسوم: إضافة وسم
لا توجد وسوم, كن أول من يضع وسما على هذه التسجيلة!
_version_ 1870679717417320448
author Zogheib, Bashar
author_facet Zogheib, Bashar
author_role author
dc.creator.none.fl_str_mv Zogheib, Bashar
dc.date.none.fl_str_mv 2019
2020-04-10T14:16:58Z
2020-04-10T14:16:58Z
dc.identifier.none.fl_str_mv https://doi.org/10.1080/24709360.2019.1615770
https://dspace.auk.edu.kw/handle/11675/5752
dc.publisher.none.fl_str_mv Taylor & Francis
dc.relation.none.fl_str_mv Biostatistics and Epidemiology
dc.title.none.fl_str_mv How many clusters exist? Answer via maximum clustering similarity implemented in R
dc.type.none.fl_str_mv Journal Article
info:eu-repo/semantics/publishedVersion
description Finding the number of clusters in a data set is considered as one of the fundamental problems in cluster analysis. This paper integrates maximum clustering similarity (MCS), for finding the optimal number of clusters, into R©statistical software through the package MCSim. The similarity between the two clustering methods is calculated at the same number of clusters, using Rand [Objective criteria for the evaluation of clustering methods. J Am Stat Assoc. 1971;66:846–850.] and Jaccard [The distribution of the flora of the alpine zone. New Phytologist. 1912;11:37–50.] indices, corrected for chance agreement. The number of clusters at which the index attains its maximum with most frequency is a candidate for the optimal number of clusters. Unlike other criteria, MCS can be used with circular data. Seven clustering algorithms, existing in R©, are implemented in MCSim. A graph of the number of clusters vs. clusters similarity using corrected similarity indices is produced. Values of the similarity indices and a clustering tree (dendrogram) are produced. Several examples including simulated, real, and circular data sets are presented to show how MCSim successfully works in practice.
id AUKR_5baaa5e3fc75df6d7e80388d73f5db3b
network_acronym_str AUKR
network_name_str AU Kuwait Rep
oai_identifier_str oai:dspace.auk.edu.kw:11675/5752
publishDate 2019
publisher.none.fl_str_mv Taylor & Francis
repository.mail.fl_str_mv
repository.name.fl_str_mv
repository_id_str
spelling How many clusters exist? Answer via maximum clustering similarity implemented in RZogheib, BasharFinding the number of clusters in a data set is considered as one of the fundamental problems in cluster analysis. This paper integrates maximum clustering similarity (MCS), for finding the optimal number of clusters, into R©statistical software through the package MCSim. The similarity between the two clustering methods is calculated at the same number of clusters, using Rand [Objective criteria for the evaluation of clustering methods. J Am Stat Assoc. 1971;66:846–850.] and Jaccard [The distribution of the flora of the alpine zone. New Phytologist. 1912;11:37–50.] indices, corrected for chance agreement. The number of clusters at which the index attains its maximum with most frequency is a candidate for the optimal number of clusters. Unlike other criteria, MCS can be used with circular data. Seven clustering algorithms, existing in R©, are implemented in MCSim. A graph of the number of clusters vs. clusters similarity using corrected similarity indices is produced. Values of the similarity indices and a clustering tree (dendrogram) are produced. Several examples including simulated, real, and circular data sets are presented to show how MCSim successfully works in practice.Taylor & Francis2020-04-10T14:16:58Z2020-04-10T14:16:58Z2019Journal Articleinfo:eu-repo/semantics/publishedVersionhttps://doi.org/10.1080/24709360.2019.1615770https://dspace.auk.edu.kw/handle/11675/5752Biostatistics and Epidemiologyoai:dspace.auk.edu.kw:11675/57522022-01-13T09:22:00Z
spellingShingle How many clusters exist? Answer via maximum clustering similarity implemented in R
Zogheib, Bashar
status_str publishedVersion
title How many clusters exist? Answer via maximum clustering similarity implemented in R
title_full How many clusters exist? Answer via maximum clustering similarity implemented in R
title_fullStr How many clusters exist? Answer via maximum clustering similarity implemented in R
title_full_unstemmed How many clusters exist? Answer via maximum clustering similarity implemented in R
title_short How many clusters exist? Answer via maximum clustering similarity implemented in R
title_sort How many clusters exist? Answer via maximum clustering similarity implemented in R
url https://doi.org/10.1080/24709360.2019.1615770
https://dspace.auk.edu.kw/handle/11675/5752