IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2512.05948.html
   My bibliography  Save this paper

Developing synthetic microdata through machine learning for firm-level business surveys

Author

Listed:
  • Jorge Cisneros
  • Timothy Wojan
  • Matthew Williams
  • Jennifer Ozawa
  • Robert Chew
  • Kimberly Janda
  • Timothy Navarro
  • Michael Floyd
  • Christine Task
  • Damon Streat

Abstract

Public-use microdata samples (PUMS) from the United States (US) Census Bureau on individuals have been available for decades. However, large increases in computing power and the greater availability of Big Data have dramatically increased the probability of re-identifying anonymized data, potentially violating the pledge of confidentiality given to survey respondents. Data science tools can be used to produce synthetic data that preserve critical moments of the empirical data but do not contain the records of any existing individual respondent or business. Developing public-use firm data from surveys presents unique challenges different from demographic data, because there is a lack of anonymity and certain industries can be easily identified in each geographic area. This paper briefly describes a machine learning model used to construct a synthetic PUMS based on the Annual Business Survey (ABS) and discusses various quality metrics. Although the ABS PUMS is currently being refined and results are confidential, we present two synthetic PUMS developed for the 2007 Survey of Business Owners, similar to the ABS business data. Econometric replication of a high impact analysis published in Small Business Economics demonstrates the verisimilitude of the synthetic data to the true data and motivates discussion of possible ABS use cases.

Suggested Citation

  • Jorge Cisneros & Timothy Wojan & Matthew Williams & Jennifer Ozawa & Robert Chew & Kimberly Janda & Timothy Navarro & Michael Floyd & Christine Task & Damon Streat, 2025. "Developing synthetic microdata through machine learning for firm-level business surveys," Papers 2512.05948, arXiv.org, revised Dec 2025.
  • Handle: RePEc:arx:papers:2512.05948
    as

    Download full text from publisher

    File URL: http://arxiv.org/pdf/2512.05948
    File Function: Latest version
    Download Restriction: no
    ---><---

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2512.05948. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: http://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.