IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2606.15031.html

Partial Identification from LLM Prompts

Author

Listed:
  • Xiaohong Chen
  • Ashesh Rambachan
  • Elie Tamer

Abstract

Large language models are increasingly used as binary classifiers when the true label is latent. We study partial identification of the prevalence $\theta = P(X^* = 1)$ from panels of LLM reports whose errors may be arbitrarily dependent given the truth. The design of replication determines the observable, and hence the identifying content: repeated prompts to one model yield a count, several named models a response vector, and both a response matrix. Cast as a two-component finite mixture, the problem makes the identification failure transparent: absent restrictions that separate the latent components, the prevalence $\theta$ is completely unidentified, and weak stochastic-ordering restrictions (first-order dominance, monotone likelihood ratio, mean ordering) leave the identified set at $[0,1]$. Identifying power comes instead from externally calibrated scores and events, which discipline the mixture in the spirit of the misclassification and corrupted-data literature. We characterize the resulting bounds, establishing validity and sharpness, and give an exact account of the identifying information in the full score distribution beyond its mean. When named models are asked repeated versions of the same question, what identifies $\theta$ is not the number of positive answers but which models agree across prompts -- a feature a vote count discards. An extension derives implied bounds on regression coefficients when $X^*$ is a regressor of interest that is not directly observed.

Suggested Citation

  • Xiaohong Chen & Ashesh Rambachan & Elie Tamer, 2026. "Partial Identification from LLM Prompts," Papers 2606.15031, arXiv.org, revised Jun 2026.
  • Handle: RePEc:arx:papers:2606.15031
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2606.15031
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. A. P. Dawid & A. M. Skene, 1979. "Maximum Likelihood Estimation of Observer Error‐Rates Using the EM Algorithm," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 28(1), pages 20-28, March.
    2. Tamer, Elie, 2010. "Partial Identification in Econometrics," Scholarly Articles 34728615, Harvard University Department of Economics.
    3. Bollinger, Christopher R., 1996. "Bounding mean regressions when a binary regressor is mismeasured," Journal of Econometrics, Elsevier, vol. 73(2), pages 387-399, August.
    4. Elie Tamer, 2010. "Partial Identification in Econometrics," Annual Review of Economics, Annual Reviews, vol. 2(1), pages 167-195, September.
    5. Marc Henry & Yuichi Kitamura & Bernard Salanié, 2014. "Partial identification of finite mixtures in econometric models," Quantitative Economics, Econometric Society, vol. 5, pages 123-144, March.
    6. Molinari, Francesca, 2008. "Partial identification of probability distributions with misclassified data," Journal of Econometrics, Elsevier, vol. 144(1), pages 81-117, May.
    7. Horowitz, Joel L & Manski, Charles F, 1995. "Identification and Robustness with Contaminated and Corrupted Data," Econometrica, Econometric Society, vol. 63(2), pages 281-302, March.
    8. Hu, Yingyao, 2008. "Identification and estimation of nonlinear models with misclassification error using instrumental variables: A general solution," Journal of Econometrics, Elsevier, vol. 144(1), pages 27-61, May.
    9. Aprajit Mahajan, 2006. "Identification and Estimation of Regression Models with Misclassification," Econometrica, Econometric Society, vol. 74(3), pages 631-665, May.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Giovanni Compiani & Yuichi Kitamura, 2016. "Using mixtures in econometric models: a brief review and some new results," Econometrics Journal, Royal Economic Society, vol. 19(3), pages 95-127, October.
    2. Takahide Yanagi, 2019. "Inference on local average treatment effects for misclassified treatment," Econometric Reviews, Taylor & Francis Journals, vol. 38(8), pages 938-960, September.
    3. Acerenza, Santiago & Ban, Kyunghoon & Kedagni, Desire, 2021. "Marginal Treatment Effects with Misclassified Treatment," ISU General Staff Papers 202106180700001132, Iowa State University, Department of Economics.
    4. Schennach, Susanne M., 2020. "Mismeasured and unobserved variables," Handbook of Econometrics, in: Steven N. Durlauf & Lars Peter Hansen & James J. Heckman & Rosa L. Matzkin (ed.), Handbook of Econometrics, edition 1, volume 7, chapter 0, pages 487-565, Elsevier.
    5. DiTraglia, Francis J. & García-Jimeno, Camilo, 2019. "Identifying the effect of a mis-classified, binary, endogenous regressor," Journal of Econometrics, Elsevier, vol. 209(2), pages 376-390.
    6. Francis J. DiTraglia & Camilo García-Jimeno, 2017. "Mis-classified, Binary, Endogenous Regressors: Identification and Inference," NBER Working Papers 23814, National Bureau of Economic Research, Inc.
    7. Hu, Yingyao, 2017. "The Econometrics of Unobservables -- Latent Variable and Measurement Error Models and Their Applications in Empirical Industrial Organization and Labor Economics [The Econometrics of Unobservables]," Economics Working Paper Archive 64578, The Johns Hopkins University,Department of Economics, revised 2021.
    8. Denni Tommasi & Arthur Lewbel & Rossella Calvi, 2017. "LATE with Mismeasured or Misspecified Treatment: An application to Women's Empowerment in India," Working Papers ECARES ECARES 2017-27, ULB -- Universite Libre de Bruxelles.
    9. Hu, Yingyao, 2017. "The econometrics of unobservables: Applications of measurement error models in empirical industrial organization and labor economics," Journal of Econometrics, Elsevier, vol. 200(2), pages 154-168.
    10. Arthur Lewbel & Xi Qu & Xun Tang, 2024. "Estimating Social Network Models with Link Misclassification," Boston College Working Papers in Economics 1079, Boston College Department of Economics.
    11. Akanksha Negi & Digvijay S. Negi, 2025. "Difference‐in‐Differences With a Misclassified Treatment," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 40(4), pages 411-423, June.
    12. Francis DiTraglia & Camilo Garcia-Jimeno, 2015. "On Mis-measured Binary Regressors: New Results And Some Comments on the Literature, Third Version," PIER Working Paper Archive 15-040, Penn Institute for Economic Research, Department of Economics, University of Pennsylvania, revised 24 Nov 2015.
    13. Battistin, Erich & De Nadai, Michele & Vuri, Daniela, 2017. "Counting rotten apples: Student achievement and score manipulation in Italian elementary Schools," Journal of Econometrics, Elsevier, vol. 200(2), pages 344-362.
    14. Lin, Zhongjian & Hu, Yingyao, 2024. "Binary choice with misclassification and social interactions, with an application to peer effects in attitude," Journal of Econometrics, Elsevier, vol. 238(1).
    15. Kline, Brendan, 2015. "Identification of complete information games," Journal of Econometrics, Elsevier, vol. 189(1), pages 117-131.
    16. Tommasi, Denni & Zhang, Lina, 2024. "Bounding program benefits when participation is misreported," Journal of Econometrics, Elsevier, vol. 238(1).
    17. Molinari, Francesca, 2020. "Microeconometrics with partial identification," Handbook of Econometrics, in: Steven N. Durlauf & Lars Peter Hansen & James J. Heckman & Rosa L. Matzkin (ed.), Handbook of Econometrics, edition 1, volume 7, chapter 0, pages 355-486, Elsevier.
    18. Meyer, Bruce D. & Mittag, Nikolas, 2017. "Misclassification in binary choice models," Journal of Econometrics, Elsevier, vol. 200(2), pages 295-311.
    19. Denni Tommasi & Lina Zhang, 2024. "Identifying program benefits when participation is misreported," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 39(6), pages 1123-1148, September.
    20. Kyunghoon Ban & D'esir'e K'edagni & Santiago Acerenza, 2021. "Local Average and Marginal Treatment Effects with a Misclassified Treatment," Papers 2105.00358, arXiv.org, revised Jun 2026.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2606.15031. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.