IDEAS home Printed from https://ideas.repec.org/a/bpj/sagmbi/v7y2008i1n24.html
   My bibliography  Save this article

Estimating Number of Clusters Based on a General Similarity Matrix with Application to Microarray Data

Author

Listed:
  • Fallah Shafagh

    (University of Toronto)

  • Tritchler David

    (University Health Network, Toronto; University of Toronto; and SUNY at Buffalo)

  • Beyene Joseph

    (Hospital for Sick Children Research Institute and University of Toronto)

Abstract

Many clustering methods require that the number of clusters believed present in a given data set be specified a priori, and a number of methods for estimating the number of clusters have been developed. However, the selection of the number of clusters is well recognized as a difficult and open problem and there is a need for methods which can shed light on specific aspects of the data. This paper adopts a model for clustering based on a specific structure for a similarity matrix. Publicly available gene expression data sets are analyzed to illustrate the method and the performance of our method is assessed by simulation.

Suggested Citation

  • Fallah Shafagh & Tritchler David & Beyene Joseph, 2008. "Estimating Number of Clusters Based on a General Similarity Matrix with Application to Microarray Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 7(1), pages 1-25, August.
  • Handle: RePEc:bpj:sagmbi:v:7:y:2008:i:1:n:24
    DOI: 10.2202/1544-6115.1261
    as

    Download full text from publisher

    File URL: https://doi.org/10.2202/1544-6115.1261
    Download Restriction: For access to full text, subscription to the journal or payment for the individual article is required.

    File URL: https://libkey.io/10.2202/1544-6115.1261?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Robert Tibshirani & Guenther Walther & Trevor Hastie, 2001. "Estimating the number of clusters in a data set via the gap statistic," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 63(2), pages 411-423.
    2. Li, Baibing & Martin, Elaine B. & Morris, A. Julian, 2002. "On principal component analysis in L1," Computational Statistics & Data Analysis, Elsevier, vol. 40(3), pages 471-474, September.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Tao Li & Yi Zhang & Dingding Wang & Jian Xu, 2019. "MCC: a Multiple Consensus Clustering Framework," Journal of Classification, Springer;The Classification Society, vol. 36(3), pages 414-434, October.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Komarek, Adam M. & Kwon, Hoyoung & Haile, Beliyou & Thierfelder, Christian & Mutenje, Munyaradzi J. & Azzarri, Carlo, 2019. "From plot to scale: ex-ante assessment of conservation agriculture in Zambia," Agricultural Systems, Elsevier, vol. 173(C), pages 504-518.
    2. Seoung Bum Kim & Jung Woo Lee & Sin Young Kim & Deok Won Lee, 2013. "Dental Informatics to Characterize Patients with Dentofacial Deformities," PLOS ONE, Public Library of Science, vol. 8(8), pages 1-8, August.
    3. Edoardo Saccenti & Johan A Westerhuis & Age K Smilde & Mariët J van der Werf & Jos A Hageman & Margriet M W B Hendriks, 2011. "Simplivariate Models: Uncovering the Underlying Biology in Functional Genomics Data," PLOS ONE, Public Library of Science, vol. 6(6), pages 1-13, June.
    4. Juan Carlos Chávez & Felipe J. Fonseca & Manuel Gómez-Zaldívar, 2017. "Resoluciones de disputas comerciales y desempeño económico regional en México. (Commercial Disputes Resolution and Regional Economic Performance in Mexico)," Ensayos Revista de Economia, Universidad Autonoma de Nuevo Leon, Facultad de Economia, vol. 0(1), pages 79-93, May.
    5. Chen, Ray-Bing & Chen, Ying & Härdle, Wolfgang K., 2014. "TVICA—Time varying independent component analysis and its application to financial data," Computational Statistics & Data Analysis, Elsevier, vol. 74(C), pages 95-109.
    6. Yan Yu Chen & Chun-Cheih Chao & Fu-Chen Liu & Po-Chen Hsu & Hsueh-Fen Chen & Shih-Chi Peng & Yung-Jen Chuang & Chung-Yu Lan & Wen-Ping Hsieh & David Shan Hill Wong, 2013. "Dynamic Transcript Profiling of Candida albicans Infection in Zebrafish: A Pathogen-Host Interaction Study," PLOS ONE, Public Library of Science, vol. 8(9), pages 1-16, September.
    7. Plat, Richard, 2009. "Stochastic portfolio specific mortality and the quantification of mortality basis risk," Insurance: Mathematics and Economics, Elsevier, vol. 45(1), pages 123-132, August.
    8. Kondylis, Athanassios & Whittaker, Joe, 2008. "Spectral preconditioning of Krylov spaces: Combining PLS and PC regression," Computational Statistics & Data Analysis, Elsevier, vol. 52(5), pages 2588-2603, January.
    9. Simplice A. Asongu & Nicholas M. Odhiambo, 2019. "Governance, capital flight and industrialisation in Africa," Journal of Economic Structures, Springer;Pan-Pacific Association of Input-Output Studies (PAPAIOS), vol. 8(1), pages 1-22, December.
    10. Thiemo Fetzer & Samuel Marden, 2017. "Take What You Can: Property Rights, Contestability and Conflict," Economic Journal, Royal Economic Society, vol. 0(601), pages 757-783, May.
    11. M. J. Aziakpono & S. Kleimeier & H. Sander, 2012. "Banking market integration in the SADC countries: evidence from interest rate analyses," Applied Economics, Taylor & Francis Journals, vol. 44(29), pages 3857-3876, October.
    12. Bianca Maria Colosimo & Luca Pagani & Marco Grasso, 2024. "Modeling spatial point processes in video-imaging via Ripley’s K-function: an application to spatter analysis in additive manufacturing," Journal of Intelligent Manufacturing, Springer, vol. 35(1), pages 429-447, January.
    13. Daniel Agness & Travis Baseler & Sylvain Chassang & Pascaline Dupas & Erik Snowberg, 2022. "Valuing the Time of the Self-Employed," Working Papers 2022-2, Princeton University. Economics Department..
    14. Batool, Fatima & Hennig, Christian, 2021. "Clustering with the Average Silhouette Width," Computational Statistics & Data Analysis, Elsevier, vol. 158(C).
    15. Ouyang, Yaofu & Li, Peng, 2018. "On the nexus of financial development, economic growth, and energy consumption in China: New perspective from a GMM panel VAR approach," Energy Economics, Elsevier, vol. 71(C), pages 238-252.
    16. Fan, Cheng & Sun, Yongjun & Zhao, Yang & Song, Mengjie & Wang, Jiayuan, 2019. "Deep learning-based feature engineering methods for improved building energy prediction," Applied Energy, Elsevier, vol. 240(C), pages 35-45.
    17. Ionela Munteanu & Adriana Grigorescu & Elena Condrea & Elena Pelinescu, 2020. "Convergent Insights for Sustainable Development and Ethical Cohesion: An Empirical Study on Corporate Governance in Romanian Public Entities," Sustainability, MDPI, vol. 12(7), pages 1-17, April.
    18. Daniel Boss & Annick Hoffmann & Benjamin Rappaz & Christian Depeursinge & Pierre J Magistretti & Dimitri Van de Ville & Pierre Marquet, 2012. "Spatially-Resolved Eigenmode Decomposition of Red Blood Cells Membrane Fluctuations Questions the Role of ATP in Flickering," PLOS ONE, Public Library of Science, vol. 7(8), pages 1-10, August.
    19. Doukas, Haris & Papadopoulou, Alexandra & Savvakis, Nikolaos & Tsoutsos, Theocharis & Psarras, John, 2012. "Assessing energy sustainability of rural communities using Principal Component Analysis," Renewable and Sustainable Energy Reviews, Elsevier, vol. 16(4), pages 1949-1957.
    20. Nicoleta Serban & Huijing Jiang, 2012. "Multilevel Functional Clustering Analysis," Biometrics, The International Biometric Society, vol. 68(3), pages 805-814, September.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bpj:sagmbi:v:7:y:2008:i:1:n:24. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Peter Golla (email available below). General contact details of provider: https://www.degruyter.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.