IDEAS home Printed from https://ideas.repec.org/a/spr/jclass/v40y2023i2d10.1007_s00357-023-09437-z.html
   My bibliography  Save this article

Zero-Inflated Time Series Clustering Via Ensemble Thick-Pen Transform

Author

Listed:
  • Minji Kim

    (University of North Carolina at Chapel Hill)

  • Hee-Seok Oh

    (Seoul National University)

  • Yaeji Lim

    (Chung-Ang University)

Abstract

This study develops a new clustering method for high-dimensional zero-inflated time series data. The proposed method is based on thick-pen transform (TPT), in which the basic idea is to draw along the data with a pen of a given thickness. Since TPT is a multi-scale visualization technique, it provides some information on the temporal tendency of neighborhood values. We introduce a modified TPT, termed ‘ensemble TPT (e-TPT)’, to enhance the temporal resolution of zero-inflated time series data that is crucial for clustering them efficiently. Furthermore, this study defines a modified similarity measure for zero-inflated time series data considering e-TPT and proposes an efficient iterative clustering algorithm suitable for the proposed measure. Finally, the effectiveness of the proposed method is demonstrated by simulation experiments and two real datasets: step count data and newly confirmed COVID-19 case data.

Suggested Citation

  • Minji Kim & Hee-Seok Oh & Yaeji Lim, 2023. "Zero-Inflated Time Series Clustering Via Ensemble Thick-Pen Transform," Journal of Classification, Springer;The Classification Society, vol. 40(2), pages 407-431, July.
  • Handle: RePEc:spr:jclass:v:40:y:2023:i:2:d:10.1007_s00357-023-09437-z
    DOI: 10.1007/s00357-023-09437-z
    as

    Download full text from publisher

    File URL: http://link.springer.com/10.1007/s00357-023-09437-z
    File Function: Abstract
    Download Restriction: Access to the full text of the articles in this series is restricted.

    File URL: https://libkey.io/10.1007/s00357-023-09437-z?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Fryzlewicz, Piotr & Oh, H. S., 2011. "Thick pen transformation for time series," LSE Research Online Documents on Economics 37663, London School of Economics and Political Science, LSE Library.
    2. P. Fryzlewicz & H.‐S. Oh, 2011. "Thick pen transformation for time series," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 73(4), pages 499-529, September.
    3. Leisch, Friedrich, 2006. "A toolbox for K-centroids cluster analysis," Computational Statistics & Data Analysis, Elsevier, vol. 51(2), pages 526-544, November.
    4. Lim, Hwa Kyung & Li, Wai Keung & Yu, Philip L.H., 2014. "Zero-inflated Poisson regression mixture model," Computational Statistics & Data Analysis, Elsevier, vol. 71(C), pages 151-158.
    5. Robert Tibshirani & Guenther Walther & Trevor Hastie, 2001. "Estimating the number of clusters in a data set via the gap statistic," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 63(2), pages 411-423.
    6. Fryzlewicz, Piotr & Ombao, Hernando, 2009. "Consistent Classification of Nonstationary Time Series Using Stochastic Wavelet Representations," Journal of the American Statistical Association, American Statistical Association, vol. 104(485), pages 299-312.
    7. Lawrence Hubert & Phipps Arabie, 1985. "Comparing partitions," Journal of Classification, Springer;The Classification Society, vol. 2(1), pages 193-218, December.
    8. Fryzlewicz, Piotr & Ombao, Hernando, 2009. "Consistent classification of non-stationary time series using stochastic wavelet representations," LSE Research Online Documents on Economics 25162, London School of Economics and Political Science, LSE Library.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Grn, Bettina & Leisch, Friedrich, 2009. "Dealing with label switching in mixture models under genuine multimodality," Journal of Multivariate Analysis, Elsevier, vol. 100(5), pages 851-861, May.
    2. Rainer Dangl & Friedrich Leisch, 2020. "Effects of Resampling in Determining the Number of Clusters in a Data Set," Journal of Classification, Springer;The Classification Society, vol. 37(3), pages 558-583, October.
    3. Li, Pai-Ling & Chiou, Jeng-Min, 2011. "Identifying cluster number for subspace projected functional data clustering," Computational Statistics & Data Analysis, Elsevier, vol. 55(6), pages 2090-2103, June.
    4. Yaeji Lim & Hee-Seok Oh & Ying Kuen Cheung, 2019. "Multiscale Clustering for Functional Data," Journal of Classification, Springer;The Classification Society, vol. 36(2), pages 368-391, July.
    5. J. Fernando Vera & Rodrigo Macías, 2021. "On the Behaviour of K-Means Clustering of a Dissimilarity Matrix by Means of Full Multidimensional Scaling," Psychometrika, Springer;The Psychometric Society, vol. 86(2), pages 489-513, June.
    6. Boztug, Yasemin & Reutterer, Thomas, 2008. "A combined approach for segment-specific market basket analysis," European Journal of Operational Research, Elsevier, vol. 187(1), pages 294-312, May.
    7. Zhiguang Huo & Li Zhu & Tianzhou Ma & Hongcheng Liu & Song Han & Daiqing Liao & Jinying Zhao & George Tseng, 2020. "Two-Way Horizontal and Vertical Omics Integration for Disease Subtype Discovery," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 12(1), pages 1-22, April.
    8. Francesco Dotto & Alessio Farcomeni & Luis Angel García-Escudero & Agustín Mayo-Iscar, 2017. "A fuzzy approach to robust regression clustering," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 11(4), pages 691-710, December.
    9. Nilsen Gro & LiestØl Knut & Borgan Ørnulf & Lingjærde Ole Christian, 2013. "Identifying clusters in genomics data by recursive partitioning," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 12(5), pages 637-652, October.
    10. Weinand, J.M. & McKenna, R. & Fichtner, W., 2019. "Developing a municipality typology for modelling decentralised energy systems," Utilities Policy, Elsevier, vol. 57(C), pages 75-96.
    11. Seoung Bum Kim & Jung Woo Lee & Sin Young Kim & Deok Won Lee, 2013. "Dental Informatics to Characterize Patients with Dentofacial Deformities," PLOS ONE, Public Library of Science, vol. 8(8), pages 1-8, August.
    12. Ja‐Yoon Jang & Hee‐Seok Oh & Yaeji Lim & Ying Kuen Cheung, 2021. "Ensemble clustering for step data via binning," Biometrics, The International Biometric Society, vol. 77(1), pages 293-304, March.
    13. Sebastian Krey & Uwe Ligges & Friedrich Leisch, 2014. "Music and timbre segmentation by recursive constrained K-means clustering," Computational Statistics, Springer, vol. 29(1), pages 37-50, February.
    14. Douglas Steinley, 2007. "Validating Clusters with the Lower Bound for Sum-of-Squares Error," Psychometrika, Springer;The Psychometric Society, vol. 72(1), pages 93-106, March.
    15. Jonathon J. O’Brien & Michael T. Lawson & Devin K. Schweppe & Bahjat F. Qaqish, 2020. "Suboptimal Comparison of Partitions," Journal of Classification, Springer;The Classification Society, vol. 37(2), pages 435-461, July.
    16. Jach, Agnieszka, 2017. "International stock market comovement in time and scale outlined with a thick pen," Journal of Empirical Finance, Elsevier, vol. 43(C), pages 115-129.
    17. Lingsong Meng & Dorina Avram & George Tseng & Zhiguang Huo, 2022. "Outcome‐guided sparse K‐means for disease subtype discovery via integrating phenotypic data with high‐dimensional transcriptomic data," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 71(2), pages 352-375, March.
    18. Xuehai Wang & Michael Nissen & Deanne Gracias & Manabu Kusakabe & Guillermo Simkin & Aixiang Jiang & Gerben Duns & Clementine Sarkozy & Laura Hilton & Elizabeth A. Chavez & Gabriela C. Segat & Rachel , 2022. "Single-cell profiling reveals a memory B cell-like subtype of follicular lymphoma with increased transformation risk," Nature Communications, Nature, vol. 13(1), pages 1-15, December.
    19. Yasemin Boztug & Thomas Reutterer, 2006. "A Combined Approach for Segment-Specific Analysis of Market Basket Data," SFB 649 Discussion Papers SFB649DP2006-006, Sonderforschungsbereich 649, Humboldt University, Berlin, Germany.
    20. Clémençon, Stéphan, 2014. "A statistical view of clustering performance through the theory of U-processes," Journal of Multivariate Analysis, Elsevier, vol. 124(C), pages 42-56.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:jclass:v:40:y:2023:i:2:d:10.1007_s00357-023-09437-z. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.