IDEAS home Printed from https://ideas.repec.org/a/plo/pcbi00/1009279.html
   My bibliography  Save this article

Eliminating accidental deviations to minimize generalization error and maximize replicability: Applications in connectomics and genomics

Author

Listed:
  • Eric W Bridgeford
  • Shangsi Wang
  • Zeyi Wang
  • Ting Xu
  • Cameron Craddock
  • Jayanta Dey
  • Gregory Kiar
  • William Gray-Roncal
  • Carlo Colantuoni
  • Christopher Douville
  • Stephanie Noble
  • Carey E Priebe
  • Brian Caffo
  • Michael Milham
  • Xi-Nian Zuo
  • Consortium for Reliability and Reproducibility
  • Joshua T Vogelstein

Abstract

Replicability, the ability to replicate scientific findings, is a prerequisite for scientific discovery and clinical utility. Troublingly, we are in the midst of a replicability crisis. A key to replicability is that multiple measurements of the same item (e.g., experimental sample or clinical participant) under fixed experimental constraints are relatively similar to one another. Thus, statistics that quantify the relative contributions of accidental deviations—such as measurement error—as compared to systematic deviations—such as individual differences—are critical. We demonstrate that existing replicability statistics, such as intra-class correlation coefficient and fingerprinting, fail to adequately differentiate between accidental and systematic deviations in very simple settings. We therefore propose a novel statistic, discriminability, which quantifies the degree to which an individual’s samples are relatively similar to one another, without restricting the data to be univariate, Gaussian, or even Euclidean. Using this statistic, we introduce the possibility of optimizing experimental design via increasing discriminability and prove that optimizing discriminability improves performance bounds in subsequent inference tasks. In extensive simulated and real datasets (focusing on brain imaging and demonstrating on genomics), only optimizing data discriminability improves performance on all subsequent inference tasks for each dataset. We therefore suggest that designing experiments and analyses to optimize discriminability may be a crucial step in solving the replicability crisis, and more generally, mitigating accidental measurement error.Author summary: In recent decades, the size and complexity of data has grown exponentially. Unfortunately, the increased scale of modern datasets brings many new challenges. At present, we are in the midst of a replicability crisis, in which scientific discoveries fail to replicate to new datasets. Difficulties in the measurement procedure and measurement processing pipelines coupled with the influx of complex high-resolution measurements, we believe, are at the core of the replicability crisis. If measurements themselves are not replicable, what hope can we have that we will be able to use the measurements for replicable scientific findings? We introduce the “discriminability” statistic, which quantifies how discriminable measurements are from one another, without limitations on the structure of the underlying measurements. We prove that discriminable strategies tend to be strategies which provide better accuracy on downstream scientific questions. We demonstrate the utility of discriminability over competing approaches in this context on two disparate datasets from both neuroimaging and genomics. Together, we believe these results suggest the value of designing experimental protocols and analysis procedures which optimize the discriminability.

Suggested Citation

  • Eric W Bridgeford & Shangsi Wang & Zeyi Wang & Ting Xu & Cameron Craddock & Jayanta Dey & Gregory Kiar & William Gray-Roncal & Carlo Colantuoni & Christopher Douville & Stephanie Noble & Carey E Prieb, 2021. "Eliminating accidental deviations to minimize generalization error and maximize replicability: Applications in connectomics and genomics," PLOS Computational Biology, Public Library of Science, vol. 17(9), pages 1-20, September.
  • Handle: RePEc:plo:pcbi00:1009279
    DOI: 10.1371/journal.pcbi.1009279
    as

    Download full text from publisher

    File URL: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1009279
    Download Restriction: no

    File URL: https://journals.plos.org/ploscompbiol/article/file?id=10.1371/journal.pcbi.1009279&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pcbi.1009279?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. John P A Ioannidis, 2005. "Why Most Published Research Findings Are False," PLOS Medicine, Public Library of Science, vol. 2(8), pages 1-1, August.
    2. Vogelstein, Joshua T., 2020. "P-Values in a Post-Truth World," OSF Preprints yw6sr, Center for Open Science.
    3. Zeileis, Achim, 2006. "Object-oriented Computation of Sandwich Estimators," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 16(i09).
    4. Nathan W Churchill & Robyn Spring & Babak Afshin-Pour & Fan Dong & Stephen C Strother, 2015. "An Automated, Adaptive Framework for Optimizing Preprocessing Pipelines in Task-Based Functional MRI," PLOS ONE, Public Library of Science, vol. 10(7), pages 1-25, July.
    5. Xi-Nian Zuo & Ting Xu & Michael Peter Milham, 2019. "Harnessing reliability for neuroscience research," Nature Human Behaviour, Nature, vol. 3(8), pages 768-771, August.
    6. Ronald D. Fricker & Katherine Burke & Xiaoyan Han & William H. Woodall, 2019. "Assessing the Statistical Analyses Used in Basic and Applied Social Psychology After Their p-Value Ban," The American Statistician, Taylor & Francis Journals, vol. 73(S1), pages 374-384, March.
    7. Jeffrey T. Leek & Roger D. Peng, 2015. "Statistics: P values are just the tip of the iceberg," Nature, Nature, vol. 520(7549), pages 612-612, April.
    8. Berna Devezer & Luis G Nardin & Bert Baumgaertner & Erkan Ozge Buzbas, 2019. "Scientific discovery in a model-centric framework: Reproducibility, innovation, and epistemic diversity," PLOS ONE, Public Library of Science, vol. 14(5), pages 1-23, May.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Uwe Hassler & Marc‐Oliver Pohle, 2022. "Unlucky Number 13? Manipulating Evidence Subject to Snooping," International Statistical Review, International Statistical Institute, vol. 90(2), pages 397-410, August.
    2. Keith R Lohse & Kristin L Sainani & J Andrew Taylor & Michael L Butson & Emma J Knight & Andrew J Vickers, 2020. "Systematic review of the use of “magnitude-based inference” in sports science and medicine," PLOS ONE, Public Library of Science, vol. 15(6), pages 1-22, June.
    3. Alexander Frankel & Maximilian Kasy, 2022. "Which Findings Should Be Published?," American Economic Journal: Microeconomics, American Economic Association, vol. 14(1), pages 1-38, February.
    4. Jyotirmoy Sarkar, 2018. "Will P†Value Triumph over Abuses and Attacks?," Biostatistics and Biometrics Open Access Journal, Juniper Publishers Inc., vol. 7(4), pages 66-71, July.
    5. Stanley, T. D. & Doucouliagos, Chris, 2019. "Practical Significance, Meta-Analysis and the Credibility of Economics," IZA Discussion Papers 12458, Institute of Labor Economics (IZA).
    6. Karin Langenkamp & Bodo Rödel & Kerstin Taufenbach & Meike Weiland, 2018. "Open Access in Vocational Education and Training Research," Publications, MDPI, vol. 6(3), pages 1-12, July.
    7. Jiang, Xianfeng & Packer, Frank, 2019. "Credit ratings of Chinese firms by domestic and global agencies: Assessing the determinants and impact," Journal of Banking & Finance, Elsevier, vol. 105(C), pages 178-193.
    8. Kevin J. Boyle & Mark Morrison & Darla Hatton MacDonald & Roderick Duncan & John Rose, 2016. "Investigating Internet and Mail Implementation of Stated-Preference Surveys While Controlling for Differences in Sample Frames," Environmental & Resource Economics, Springer;European Association of Environmental and Resource Economists, vol. 64(3), pages 401-419, July.
    9. Jelte M Wicherts & Marjan Bakker & Dylan Molenaar, 2011. "Willingness to Share Research Data Is Related to the Strength of the Evidence and the Quality of Reporting of Statistical Results," PLOS ONE, Public Library of Science, vol. 6(11), pages 1-7, November.
    10. Valentine, Kathrene D & Buchanan, Erin Michelle & Scofield, John E. & Beauchamp, Marshall T., 2017. "Beyond p-values: Utilizing Multiple Estimates to Evaluate Evidence," OSF Preprints 9hp7y, Center for Open Science.
    11. Anton, Roman, 2014. "Sustainable Intrapreneurship - The GSI Concept and Strategy - Unfolding Competitive Advantage via Fair Entrepreneurship," MPRA Paper 69713, University Library of Munich, Germany, revised 01 Feb 2015.
    12. Dudek, Thomas & Brenøe, Anne Ardila & Feld, Jan & Rohrer, Julia, 2022. "No Evidence That Siblings' Gender Affects Personality across Nine Countries," IZA Discussion Papers 15137, Institute of Labor Economics (IZA).
    13. Frederique Bordignon, 2020. "Self-correction of science: a comparative study of negative citations and post-publication peer review," Scientometrics, Springer;Akadémiai Kiadó, vol. 124(2), pages 1225-1239, August.
    14. Stefan Seifert & Christoph Kahle & Silke Hüttel, 2021. "Price Dispersion in Farmland Markets: What Is the Role of Asymmetric Information?," American Journal of Agricultural Economics, John Wiley & Sons, vol. 103(4), pages 1545-1568, August.
    15. Omar Al-Ubaydli & John A. List, 2015. "Do Natural Field Experiments Afford Researchers More or Less Control than Laboratory Experiments? A Simple Model," NBER Working Papers 20877, National Bureau of Economic Research, Inc.
    16. Alice Hengevoss, 2021. "Assessing the Impact of Nonprofit Organizations on Multi-Actor Global Governance Initiatives: The Case of the UN Global Compact," Sustainability, MDPI, vol. 13(13), pages 1-13, June.
    17. Guan, Sihai & Wan, Dongyu & Yang, Yanmiao & Biswal, Bharat, 2022. "Sources of multifractality of the brain rs-fMRI signal," Chaos, Solitons & Fractals, Elsevier, vol. 160(C).
    18. Aurelie Seguin & Wolfgang Forstmeier, 2012. "No Band Color Effects on Male Courtship Rate or Body Mass in the Zebra Finch: Four Experiments and a Meta-Analysis," PLOS ONE, Public Library of Science, vol. 7(6), pages 1-11, June.
    19. Sviták, Jan & Tichem, Jan & Haasbeek, Stefan, 2021. "Price effects of search advertising restrictions," International Journal of Industrial Organization, Elsevier, vol. 77(C).
    20. Ankur Moitra & Dhruv Rohatgi, 2022. "Provably Auditing Ordinary Least Squares in Low Dimensions," Papers 2205.14284, arXiv.org, revised Jun 2022.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1009279. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.