IDEAS home Printed from https://ideas.repec.org/a/nas/journl/v122y2025pe2427298122.html
   My bibliography  Save this article

Estimating wage disparities using foundation models

Author

Listed:
  • Keyon Vafa

    (a Harvard Data Science Initiative , Harvard University , Cambridge , MA 02138)

  • Susan Athey

    (c Stanford Institute for Human-Centered Artificial Intelligence , Stanford University , Stanford , CA 94305)

  • David M. Blei

    (e Department of Statistics , Columbia University , New York , NY 10027)

Abstract

The rise of foundation models marks a paradigm shift in machine learning: instead of training specialized models from scratch, foundation models are trained on massive datasets before being adjusted or fine-tuned to make predictions on smaller datasets. Initially developed for text, foundation models have also excelled at making predictions about social science data. However, while many estimation problems in the social sciences use prediction as an intermediate step, they ultimately require different criteria for success. In this paper, we develop methods for fine-tuning foundation models to perform these estimation problems. We first characterize an omitted variable bias that can arise when a foundation model is fine-tuned in the standard way: to minimize predictive error. We then provide a set of conditions for fine-tuning under which estimates derived from a foundation model are n -consistent. Based on this theory, we develop fine-tuning algorithms that empirically mitigate this omitted variable bias. To demonstrate our ideas, we study gender wage gap estimation. Classical methods for estimating the adjusted wage gap employ simple predictive models of wages, which can induce omitted variable bias because they condition on coarse summaries of career history. Instead, we use a custom-built foundation model, capturing a richer representation of career history. Using data from the Panel Study of Income Dynamics, we find that career history explains more of the gender wage gap than standard econometric models can measure, and we identify elements of career history that are omitted by standard models but are important for explaining the gap.

Suggested Citation

  • Keyon Vafa & Susan Athey & David M. Blei, 2025. "Estimating wage disparities using foundation models," Proceedings of the National Academy of Sciences, Proceedings of the National Academy of Sciences, vol. 122(22), pages 2427298122-, June.
  • Handle: RePEc:nas:journl:v:122:y:2025:p:e2427298122
    DOI: 10.1073/pnas.2427298122
    as

    Download full text from publisher

    File URL: https://doi.org/10.1073/pnas.2427298122
    Download Restriction: no

    File URL: https://libkey.io/10.1073/pnas.2427298122?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:nas:journl:v:122:y:2025:p:e2427298122. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: PNAS Product Team (email available below). General contact details of provider: http://www.pnas.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.