Machine learning models trained on synthetic datasets of multiple sample sizes for the use of predicting blood pressure from clinical data in a national dataset

Machine learning models trained on synthetic datasets of multiple sample sizes for the use of predicting blood pressure from clinical data in a national dataset

Author

Listed:

Anmol Arora
Ananya Arora

Abstract

Introduction: The potential for synthetic data to act as a replacement for real data in research has attracted attention in recent months due to the prospect of increasing access to data and overcoming data privacy concerns when sharing data. The field of generative artificial intelligence and synthetic data is still early in its development, with a research gap evidencing that synthetic data can adequately be used to train algorithms that can be used on real data. This study compares the performance of a series machine learning models trained on real data and synthetic data, based on the National Diet and Nutrition Survey (NDNS). Methods: Features identified to be potentially of relevance by directed acyclic graphs were isolated from the NDNS dataset and used to construct synthetic datasets and impute missing data. Recursive feature elimination identified only four variables needed to predict mean arterial blood pressure: age, sex, weight and height. Bayesian generalised linear regression, random forest and neural network models were constructed based on these four variables to predict blood pressure. Models were trained on the real data training set (n = 2408), a synthetic data training set (n = 2408) and larger synthetic data training set (n = 4816) and a combination of the real and synthetic data training set (n = 4816). The same test set (n = 424) was used for each model. Results: Synthetic datasets demonstrated a high degree of fidelity with the real dataset. There was no significant difference between the performance of models trained on real, synthetic or combined datasets. Mean average error across all models and all training data ranged from 8.12 To 8.33. This indicates that synthetic data was capable of training equally accurate machine learning models as real data. Discussion: Further research is needed on a variety of datasets to confirm the utility of synthetic data to replace the use of potentially identifiable patient data. There is also further urgent research needed into evidencing that synthetic data can truly protect patient privacy against adversarial attempts to re-identify real individuals from the synthetic dataset.

Suggested Citation

Anmol Arora & Ananya Arora, 2023. "Machine learning models trained on synthetic datasets of multiple sample sizes for the use of predicting blood pressure from clinical data in a national dataset," PLOS ONE, Public Library of Science, vol. 18(3), pages 1-10, March.

Handle: RePEc:plo:pone00:0283094
DOI: 10.1371/journal.pone.0283094

Download full text from publisher

References listed on IDEAS

Chao Yan & Yao Yan & Zhiyu Wan & Ziqi Zhang & Larsson Omberg & Justin Guinney & Sean D. Mooney & Bradley A. Malin, 2022. "A Multifaceted benchmarking of synthetic electronic health record generation models," Nature Communications, Nature, vol. 13(1), pages 1-18, December.
Nowok, Beata & Raab, Gillian M. & Dibben, Chris, 2016. "synthpop: Bespoke Creation of Synthetic Data in R," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 74(i11).

Full references (including those not matched with items on IDEAS)

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Dominik Bietsch & Robert Stahlbock & Stefan Voß, 2023. "Synthetic Data as a Proxy for Real-World Electronic Health Records in the Patient Length of Stay Prediction," Sustainability, MDPI, vol. 15(18), pages 1-30, September.
Qi Chang & Zhennan Yan & Mu Zhou & Hui Qu & Xiaoxiao He & Han Zhang & Lohendran Baskaran & Subhi Al’Aref & Hongsheng Li & Shaoting Zhang & Dimitris N. Metaxas, 2023. "Mining multi-center heterogeneous medical data with distributed synthetic learning," Nature Communications, Nature, vol. 14(1), pages 1-16, December.
Brandon Theodorou & Cao Xiao & Jimeng Sun, 2023. "Synthesize high-dimensional longitudinal electronic health records via hierarchical autoregressive language model," Nature Communications, Nature, vol. 14(1), pages 1-13, December.
James Jackson & Robin Mitra & Brian Francis & Iain Dove, 2022. "Using saturated count models for user‐friendly synthesis of large confidential administrative databases," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 185(4), pages 1613-1643, October.
Fabian Sven Karst & Mahei Manhai Li & Jan Marco Leimeister, 2025. "SynDEc: A Synthetic Data Ecosystem," Electronic Markets, Springer;IIM University of St. Gallen, vol. 35(1), pages 1-28, December.
Joshua Snoke & Gillian M. Raab & Beata Nowok & Chris Dibben & Aleksandra Slavkovic, 2018. "General and specific utility measures for synthetic data," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 181(3), pages 663-688, June.
Wesley J. Marrero & Mariel S. Lavieri & Jeremy B. Sussman, 2021. "Optimal cholesterol treatment plans and genetic testing strategies for cardiovascular diseases," Health Care Management Science, Springer, vol. 24(1), pages 1-25, March.
Venkitasubramanian, Kailas, 2026. "Synthesizing Tabular Microdata with Gaussian Copulas: The rsdv R Package," SocArXiv t6cne_v1, Center for Open Science.
Seungkyu Kim & Johan Lim & Donghyeon Yu, 2025. "Synthetic data generation method providing enhanced covariance matrix estimation," Computational Statistics, Springer, vol. 40(7), pages 4007-4035, September.
Asunur Cezar & Srinivasan Raghunathan & Sumit Sarkar, 2020. "Adversarial Classification: Impact of Agents’ Faking Cost on Firms and Agents," Production and Operations Management, Production and Operations Management Society, vol. 29(12), pages 2789-2807, December.
Speidel, Matthias & Drechsler, Jörg & Jolani, Shahab, 2018. "R package hmi: a convenient tool for hierarchical multiple imputation and beyond," IAB-Discussion Paper 201816, Institut für Arbeitsmarkt- und Berufsforschung (IAB), Nürnberg [Institute for Employment Research, Nuremberg, Germany].
Lau Lilleholt & Ingo Zettler & Cornelia Betsch & Robert Böhm, 2023. "Development and validation of the pandemic fatigue scale," Nature Communications, Nature, vol. 14(1), pages 1-19, December.
Schön, Peter & Heinen, Eva & Manum, Bendik, 2026. "Route-based network analysis: “Route betweenness” and “route link density” applied to a study of a proposed bridge in Trondheim," Journal of Transport Geography, Elsevier, vol. 130(C).
Stefan Wimmer & Robert Finger, 2023. "A note on synthetic data for replication purposes in agricultural economics," Journal of Agricultural Economics, Wiley Blackwell, vol. 74(1), pages 316-323, February.
Deer, Lachlan & Adler, Susanne J. & Datta, Hannes & Mizik, Natalie & Sarstedt, Marko, 2025. "Toward open science in marketing research," International Journal of Research in Marketing, Elsevier, vol. 42(1), pages 212-233.
- Deer, Lachlan & Adler, Susanne Jana & Datta, Hannes & Mizik, Natalie & Sarstedt, Marko, 2024. "Toward Open Science in Marketing Research," OSF Preprints f7a8c, Center for Open Science.
Marta Cipriani & Lorenzo Di Rocco & Maria Puopolo & Marco Alfò, 2025. "A flexible parametric approach to synthetic patients generation using health data," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 34(4), pages 639-662, September.
Gunjan Chandra & Pekka Siirtola & Satu Tamminen & Mikael J. Knip & Riitta Veijola & Juha Röning, 2022. "Impacts of Data Synthesis: A Metric for Quantifiable Data Standards and Performances," Data, MDPI, vol. 7(12), pages 1-26, December.
Severin Elvatun & Daan Knoors & Simon Brant & Christian Jonasson & Jan F Nygård, 2025. "Synthetic data as external control arms in scarce single-arm clinical trials," PLOS Digital Health, Public Library of Science, vol. 4(1), pages 1-13, January.
Daiho Uhm & Sunghae Jun, 2022. "Zero-Inflated Patent Data Analysis Using Generating Synthetic Samples," Future Internet, MDPI, vol. 14(7), pages 1-11, July.
Felix Ritchie & Jim Smith, 2019. "Confidentiality and linked data," Papers 1907.06465, arXiv.org.

More about this item

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0283094. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Machine learning models trained on synthetic datasets of multiple sample sizes for the use of predicting blood pressure from clinical data in a national dataset

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Most related items

More about this item

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data