IDEAS home Printed from https://ideas.repec.org/a/eee/csdana/v180y2023ics0167947322002481.html
   My bibliography  Save this article

Fast and fully-automated histograms for large-scale data sets

Author

Listed:
  • Zelaya Mendizábal, Valentina
  • Boullé, Marc
  • Rossi, Fabrice

Abstract

G-Enum histograms are a new fast and fully automated method for irregular histogram construction. By framing histogram construction as a density estimation problem and its automation as a model selection task, these histograms leverage the Minimum Description Length principle (MDL) to derive two different model selection criteria. Several proven theoretical results about these criteria give insights about their asymptotic behaviour and are used to speed up their optimisation. These insights, combined to a greedy search heuristic, are used to construct histograms in linearithmic time rather than the polynomial time incurred by previous works. The capabilities of the proposed MDL density estimation method are illustrated with reference to other fully automated methods in the literature, both on synthetic and large real-world data sets.

Suggested Citation

  • Zelaya Mendizábal, Valentina & Boullé, Marc & Rossi, Fabrice, 2023. "Fast and fully-automated histograms for large-scale data sets," Computational Statistics & Data Analysis, Elsevier, vol. 180(C).
  • Handle: RePEc:eee:csdana:v:180:y:2023:i:c:s0167947322002481
    DOI: 10.1016/j.csda.2022.107668
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S0167947322002481
    Download Restriction: Full text for ScienceDirect subscribers only.

    File URL: https://libkey.io/10.1016/j.csda.2022.107668?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Peter D. Grünwald, 2007. "The Minimum Description Length Principle," MIT Press Books, The MIT Press, edition 1, volume 1, number 0262072815, December.
    2. Rozenholc, Yves & Mildenberger, Thoralf & Gather, Ursula, 2010. "Combining regular and irregular histograms by penalized likelihood," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 3313-3323, December.
    3. Charles R. Harris & K. Jarrod Millman & Stéfan J. Walt & Ralf Gommers & Pauli Virtanen & David Cournapeau & Eric Wieser & Julian Taylor & Sebastian Berg & Nathaniel J. Smith & Robert Kern & Matti Picu, 2020. "Array programming with NumPy," Nature, Nature, vol. 585(7825), pages 357-362, September.
    4. Housen Li & Axel Munk & Hannes Sieling & Guenther Walther, 2020. "The essential histogram," Biometrika, Biometrika Trust, vol. 107(2), pages 347-364.
    5. Celisse, Alain & Robin, Stephane, 2008. "Nonparametric density estimation by exact leave-p-out cross-validation," Computational Statistics & Data Analysis, Elsevier, vol. 52(5), pages 2350-2368, January.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Tan Wang & L. Jeff Hong, 2023. "Large-Scale Inventory Optimization: A Recurrent Neural Networks–Inspired Simulation Approach," INFORMS Journal on Computing, INFORMS, vol. 35(1), pages 196-215, January.
    2. Léon Faure & Bastien Mollet & Wolfram Liebermeister & Jean-Loup Faulon, 2023. "A neural-mechanistic hybrid approach improving the predictive power of genome-scale metabolic models," Nature Communications, Nature, vol. 14(1), pages 1-14, December.
    3. Claudia Quinteros-Cartaya & Guillermo Solorio-Magaña & Francisco Javier Núñez-Cornú & Felipe de Jesús Escalona-Alcázar & Diana Núñez, 2023. "Microearthquakes in the Guadalajara Metropolitan Zone, Mexico: evidence from buried active faults in Tesistán Valley, Zapopan," Natural Hazards: Journal of the International Society for the Prevention and Mitigation of Natural Hazards, Springer;International Society for the Prevention and Mitigation of Natural Hazards, vol. 116(3), pages 2797-2818, April.
    4. López Pérez, Mario & Mansilla Corona, Ricardo, 2022. "Ordinal synchronization and typical states in high-frequency digital markets," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 598(C).
    5. Van Hanh Nguyen & Catherine Matias, 2014. "On Efficient Estimators of the Proportion of True Null Hypotheses in a Multiple Testing Setup," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 41(4), pages 1167-1194, December.
    6. Jessica M. Vanslambrouck & Sean B. Wilson & Ker Sin Tan & Ella Groenewegen & Rajeev Rudraraju & Jessica Neil & Kynan T. Lawlor & Sophia Mah & Michelle Scurr & Sara E. Howden & Kanta Subbarao & Melissa, 2022. "Enhanced metanephric specification to functional proximal tubule enables toxicity screening and infectious disease modelling in kidney organoids," Nature Communications, Nature, vol. 13(1), pages 1-23, December.
    7. Kiran Krishnamachari & Dylan Lu & Alexander Swift-Scott & Anuar Yeraliyev & Kayla Lee & Weitai Huang & Sim Ngak Leng & Anders Jacobsen Skanderup, 2022. "Accurate somatic variant detection using weakly supervised deep learning," Nature Communications, Nature, vol. 13(1), pages 1-8, December.
    8. Lauren L. Porter & Allen K. Kim & Swechha Rimal & Loren L. Looger & Ananya Majumdar & Brett D. Mensh & Mary R. Starich & Marie-Paule Strub, 2022. "Many dissimilar NusG protein domains switch between α-helix and β-sheet folds," Nature Communications, Nature, vol. 13(1), pages 1-12, December.
    9. Matthew Rosenblatt & Link Tejavibulya & Rongtao Jiang & Stephanie Noble & Dustin Scheinost, 2024. "Data leakage inflates prediction performance in connectome-based machine learning models," Nature Communications, Nature, vol. 15(1), pages 1-15, December.
    10. Jackie Grant & Mark Hindmarsh & Sergey E. Koposov, 2022. "The distribution of loss to future USS pensions due to the UUK cuts of April 2022," Papers 2206.06201, arXiv.org.
    11. Sayedali Shetab Boushehri & Katharina Essig & Nikolaos-Kosmas Chlis & Sylvia Herter & Marina Bacac & Fabian J. Theis & Elke Glasmacher & Carsten Marr & Fabian Schmich, 2023. "Explainable machine learning for profiling the immunological synapse and functional characterization of therapeutic antibodies," Nature Communications, Nature, vol. 14(1), pages 1-16, December.
    12. Neuwald Andrew F., 2014. "Protein domain hierarchy Gibbs sampling strategies," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 13(4), pages 1-21, August.
    13. Shukla, Mohak & Thakur, Ajay D., 2022. "An Enquiry on similarities between Renormalization Group and Auto-Encoders using Transfer Learning," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 608(P1).
    14. Khaled Akkad & David He, 2023. "A dynamic mode decomposition based deep learning technique for prognostics," Journal of Intelligent Manufacturing, Springer, vol. 34(5), pages 2207-2224, June.
    15. Romain Fournier & Zoi Tsangalidou & David Reich & Pier Francesco Palamara, 2023. "Haplotype-based inference of recent effective population size in modern and ancient DNA samples," Nature Communications, Nature, vol. 14(1), pages 1-13, December.
    16. Laura Portell & Sergi Morera & Helena Ramalhinho, 2022. "Door-to-Door Transportation Services for Reduced Mobility Population: A Descriptive Analytics of the City of Barcelona," IJERPH, MDPI, vol. 19(8), pages 1-20, April.
    17. Caroline Haimerl & Douglas A. Ruff & Marlene R. Cohen & Cristina Savin & Eero P. Simoncelli, 2023. "Targeted V1 comodulation supports task-adaptive sensory decisions," Nature Communications, Nature, vol. 14(1), pages 1-15, December.
    18. Jonas Bunsen & Matthias Finkbeiner, 2022. "An Introductory Review of Input-Output Analysis in Sustainability Sciences Including Potential Implications of Aggregation," Sustainability, MDPI, vol. 15(1), pages 1-24, December.
    19. Petros C. Lazaridis & Ioannis E. Kavvadias & Konstantinos Demertzis & Lazaros Iliadis & Lazaros K. Vasiliadis, 2023. "Interpretable Machine Learning for Assessing the Cumulative Damage of a Reinforced Concrete Frame Induced by Seismic Sequences," Sustainability, MDPI, vol. 15(17), pages 1-31, August.
    20. Matthias Wagener & Andriette Bekker & Mohammad Arashi, 2021. "Mastering the Body and Tail Shape of a Distribution," Mathematics, MDPI, vol. 9(21), pages 1-22, October.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:180:y:2023:i:c:s0167947322002481. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.