Author
Listed:
- Dave Osthus
- Alexander C Murph
- Emma E Goldberg
- Lauren J Beesley
- William M Fischer
- Nidhi Parikh
- Lauren A Castro
Abstract
Forecasting infectious disease outbreaks is hard. Forecasting emerging infectious diseases with limited historical data is even harder. In this paper, we investigate ways to improve emerging infectious disease forecasting when little pathogen-specific training data are available. Specifically, we explore two sources of information that may be available near the start of an emerging disease outbreak: synthetic data and genetic information. For this investigation, we conducted an experiment where we trained deep learning models on different combinations of real and synthetic data, both with and without genetic information, to explore how these models compare when forecasting COVID-19 cases for US states. All models are developed with an eye towards forecasting the next pandemic. We find that models trained with synthetic data have better forecast accuracy than models trained on real data alone, and models that use genetic variants have better forecast accuracy compared to those that do not. All models outperformed a baseline persistence model, a benchmark that proved challenging for many real-time COVID-19 case forecasting models, and multiple models outperformed the COVIDHub-4_week_ensemble. This paper demonstrates the value of these underutilized sources of information and provides a blueprint for forecasting future pandemics.Author summary: Forecasting emerging infectious diseases is difficult because, by definition, little or no historical data are available for the pathogen of interest. This lack of training data is a major barrier to using flexible machine learning models in emerging disease forecasting settings. In this paper, we investigate two sources of information likely to be available near the start of a future outbreak: synthetic outbreak data and viral genetic information. Using COVID-19 as a test case, we train transformer-based models on historical respiratory disease data, synthetic outbreaks, and SARS-CoV-2 variant information to forecast weekly cases in U.S. states. Models trained with synthetic data outperform those trained on historical data alone, and models trained on both real and synthetic data perform best overall. We also find that using variant-attributable cases improves forecasts relative to forecasting total cases directly. We show these models were competitive with the best real-time COVID-19 case forecasting models. Together, these results show that synthetic data and genomic surveillance are practical, underused tools for epidemic forecasting. They offer a path toward scalable forecasting systems that can be deployed early in the next pandemic, when reliable forecasts are most needed and hardest to produce.
Suggested Citation
Dave Osthus & Alexander C Murph & Emma E Goldberg & Lauren J Beesley & William M Fischer & Nidhi Parikh & Lauren A Castro, 2026.
"Leveraging synthetic and genetic data to improve epidemic forecasting,"
PLOS Computational Biology, Public Library of Science, vol. 22(8), pages 1-29, August.
Handle:
RePEc:plo:pcbi00:1014630
DOI: 10.1371/journal.pcbi.1014630
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pcbi00:1014630. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ploscompbiol (email available below). General contact details of provider: https://journals.plos.org/ploscompbiol/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.