IDEAS home Printed from https://ideas.repec.org/a/spr/compst/v35y2020i3d10.1007_s00180-020-00970-8.html
   My bibliography  Save this article

A Bayesian perspective of statistical machine learning for big data

Author

Listed:
  • Rajiv Sambasivan

    (Chennai Mathematical Institute
    University of Southampton)

  • Sourish Das

    (Chennai Mathematical Institute
    University of Southampton)

  • Sujit K. Sahu

    (Chennai Mathematical Institute
    University of Southampton)

Abstract

Statistical Machine Learning (SML) refers to a body of algorithms and methods by which computers are allowed to discover important features of input data sets which are often very large in size. The very task of feature discovery from data is essentially the meaning of the keyword ‘learning’ in SML. Theoretical justifications for the effectiveness of the SML algorithms are underpinned by sound principles from different disciplines, such as Computer Science and Statistics. The theoretical underpinnings particularly justified by statistical inference methods are together termed as statistical learning theory. This paper provides a review of SML from a Bayesian decision theoretic point of view—where we argue that many SML techniques are closely connected to making inference by using the so called Bayesian paradigm. We discuss many important SML techniques such as supervised and unsupervised learning, deep learning, online learning and Gaussian processes especially in the context of very large data sets where these are often employed. We present a dictionary which maps the key concepts of SML from Computer Science and Statistics. We illustrate the SML techniques with three moderately large data sets where we also discuss many practical implementation issues. Thus the review is especially targeted at statisticians and computer scientists who are aspiring to understand and apply SML for moderately large to big data sets.

Suggested Citation

  • Rajiv Sambasivan & Sourish Das & Sujit K. Sahu, 2020. "A Bayesian perspective of statistical machine learning for big data," Computational Statistics, Springer, vol. 35(3), pages 893-930, September.
  • Handle: RePEc:spr:compst:v:35:y:2020:i:3:d:10.1007_s00180-020-00970-8
    DOI: 10.1007/s00180-020-00970-8
    as

    Download full text from publisher

    File URL: http://link.springer.com/10.1007/s00180-020-00970-8
    File Function: Abstract
    Download Restriction: Access to the full text of the articles in this series is restricted.

    File URL: https://libkey.io/10.1007/s00180-020-00970-8?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. David M. Blei & Alp Kucukelbir & Jon D. McAuliffe, 2017. "Variational Inference: A Review for Statisticians," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 112(518), pages 859-877, April.
    2. Park, Trevor & Casella, George, 2008. "The Bayesian Lasso," Journal of the American Statistical Association, American Statistical Association, vol. 103, pages 681-686, June.
    3. Hui Zou & Trevor Hastie, 2005. "Addendum: Regularization and variable selection via the elastic net," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 67(5), pages 768-768, November.
    4. Megan L Head & Luke Holman & Rob Lanfear & Andrew T Kahn & Michael D Jennions, 2015. "The Extent and Consequences of P-Hacking in Science," PLOS Biology, Public Library of Science, vol. 13(3), pages 1-15, March.
    5. Hui Zou & Trevor Hastie, 2005. "Regularization and variable selection via the elastic net," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 67(2), pages 301-320, April.
    6. Das, Sourish & Dey, Dipak K., 2006. "On Bayesian Analysis of Generalized Linear Models Using the Jacobian Technique," The American Statistician, American Statistical Association, vol. 60, pages 264-268, August.
    7. Das, Sourish & Dey, Dipak K., 2010. "On Bayesian inference for generalized multivariate gamma distribution," Statistics & Probability Letters, Elsevier, vol. 80(19-20), pages 1492-1499, October.
    8. Sourish Das & Dipak K. Dey, 2013. "On Dynamic Generalized Linear Models with Applications," Methodology and Computing in Applied Probability, Springer, vol. 15(2), pages 407-421, June.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Christian Soize, 2023. "Probabilistic learning constrained by realizations using a weak formulation of Fourier transform of probability measures," Computational Statistics, Springer, vol. 38(4), pages 1879-1925, December.
    2. Hirofumi Michimae & Takeshi Emura, 2022. "Bayesian ridge estimators based on copula-based joint prior distributions for regression coefficients," Computational Statistics, Springer, vol. 37(5), pages 2741-2769, November.
    3. Emanuele Dolera, 2022. "Asymptotic Efficiency of Point Estimators in Bayesian Predictive Inference," Mathematics, MDPI, vol. 10(7), pages 1-27, April.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Diego Vidaurre & Concha Bielza & Pedro Larrañaga, 2013. "A Survey of L1 Regression," International Statistical Review, International Statistical Institute, vol. 81(3), pages 361-387, December.
    2. Sweata Sen & Damitri Kundu & Kiranmoy Das, 2023. "Variable selection for categorical response: a comparative study," Computational Statistics, Springer, vol. 38(2), pages 809-826, June.
    3. Li, Jiahan & Chen, Weiye, 2014. "Forecasting macroeconomic time series: LASSO-based approaches and their forecast combinations with dynamic factor models," International Journal of Forecasting, Elsevier, vol. 30(4), pages 996-1015.
    4. Mike K. P. So & Wing Ki Liu & Amanda M. Y. Chu, 2018. "Bayesian Shrinkage Estimation Of Time-Varying Covariance Matrices In Financial Time Series," Advances in Decision Sciences, Asia University, Taiwan, vol. 22(1), pages 369-404, December.
    5. Petropoulos, Fotios & Apiletti, Daniele & Assimakopoulos, Vassilios & Babai, Mohamed Zied & Barrow, Devon K. & Ben Taieb, Souhaib & Bergmeir, Christoph & Bessa, Ricardo J. & Bijak, Jakub & Boylan, Joh, 2022. "Forecasting: theory and practice," International Journal of Forecasting, Elsevier, vol. 38(3), pages 705-871.
      • Fotios Petropoulos & Daniele Apiletti & Vassilios Assimakopoulos & Mohamed Zied Babai & Devon K. Barrow & Souhaib Ben Taieb & Christoph Bergmeir & Ricardo J. Bessa & Jakub Bijak & John E. Boylan & Jet, 2020. "Forecasting: theory and practice," Papers 2012.03854, arXiv.org, revised Jan 2022.
    6. Gary Koop & Dimitris Korobilis, 2023. "Bayesian Dynamic Variable Selection In High Dimensions," International Economic Review, Department of Economics, University of Pennsylvania and Osaka University Institute of Social and Economic Research Association, vol. 64(3), pages 1047-1074, August.
    7. Subharup Guha & Rex Jung & David Dunson, 2022. "Predicting phenotypes from brain connection structure," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 71(3), pages 639-668, June.
    8. Philip D. Waggoner & Alec Macmillen, 2022. "Pursuing open-source development of predictive algorithms: the case of criminal sentencing algorithms," Journal of Computational Social Science, Springer, vol. 5(1), pages 89-109, May.
    9. Baragatti, M. & Pommeret, D., 2012. "A study of variable selection using g-prior distribution with ridge parameter," Computational Statistics & Data Analysis, Elsevier, vol. 56(6), pages 1920-1934.
    10. Nicolai Meinshausen & Peter Bühlmann, 2010. "Stability selection," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 72(4), pages 417-473, September.
    11. Feihan Lu & Yao Zheng & Harrington Cleveland & Chris Burton & David Madigan, 2018. "Bayesian hierarchical vector autoregressive models for patient-level predictive modeling," PLOS ONE, Public Library of Science, vol. 13(12), pages 1-27, December.
    12. Roberto Casarin & Fausto Corradin & Francesco Ravazzolo & Nguyen Domenico Sartore, 2020. "A Scoring Rule for Factor and Autoregressive Models Under Misspecification," Advances in Decision Sciences, Asia University, Taiwan, vol. 24(2), pages 66-103, June.
    13. Gilles Celeux & Mohammed El Anbari & Jean-Michel Marin & Christian P. Robert, 2010. "Regularization in Regression : Comparing Bayesian and Frequentist Methods in a Poorly Informative Situation," Working Papers 2010-43, Center for Research in Economics and Statistics.
    14. Shutes, Karl & Adcock, Chris, 2013. "Regularized Extended Skew-Normal Regression," MPRA Paper 58445, University Library of Munich, Germany, revised 09 Sep 2014.
    15. Daniel Felix Ahelegbey & Monica Billio & Roberto Casarin, 2016. "Sparse Graphical Vector Autoregression: A Bayesian Approach," Annals of Economics and Statistics, GENES, issue 123-124, pages 333-361.
    16. Korobilis, Dimitris, 2013. "Hierarchical shrinkage priors for dynamic regressions with many predictors," International Journal of Forecasting, Elsevier, vol. 29(1), pages 43-59.
    17. Shutes, Karl & Adcock, Chris, 2013. "Regularized Skew-Normal Regression," MPRA Paper 52217, University Library of Munich, Germany, revised 11 Dec 2013.
    18. Banerjee, Sayantan, 2022. "Horseshoe shrinkage methods for Bayesian fusion estimation," Computational Statistics & Data Analysis, Elsevier, vol. 174(C).
    19. Matthew Gentzkow & Bryan T. Kelly & Matt Taddy, 2017. "Text as Data," NBER Working Papers 23276, National Bureau of Economic Research, Inc.
    20. Cox Lwaka Tamba & Yuan-Li Ni & Yuan-Ming Zhang, 2017. "Iterative sure independence screening EM-Bayesian LASSO algorithm for multi-locus genome-wide association studies," PLOS Computational Biology, Public Library of Science, vol. 13(1), pages 1-20, January.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:compst:v:35:y:2020:i:3:d:10.1007_s00180-020-00970-8. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.