IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2606.17165.html

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Author

Listed:
  • Joel Persson
  • M{aa}rten Schultzberg
  • Sebastian Ankargren

Abstract

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.

Suggested Citation

  • Joel Persson & M{aa}rten Schultzberg & Sebastian Ankargren, 2026. "Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference," Papers 2606.17165, arXiv.org, revised Jun 2026.
  • Handle: RePEc:arx:papers:2606.17165
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2606.17165
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Susan Athey & Raj Chetty & Guido Imbens, 2020. "Using Experiments to Correct for Selection in Observational Studies," Papers 2006.09676, arXiv.org, revised May 2025.
    2. Susan Athey & Raj Chetty & Guido W. Imbens & Hyunseung Kang, 2019. "The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely," NBER Working Papers 26463, National Bureau of Economic Research, Inc.
    3. Victor Chernozhukov & Denis Chetverikov & Mert Demirer & Esther Duflo & Christian Hansen & Whitney Newey & James Robins, 2018. "Double/debiased machine learning for treatment and structural parameters," Econometrics Journal, Royal Economic Society, vol. 21(1), pages 1-68, February.
    4. Jeremy Yang & Dean Eckles & Paramveer Dhillon & Sinan Aral, 2024. "Targeting for Long-Term Outcomes," Management Science, INFORMS, vol. 70(6), pages 3841-3855, June.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Ruoxuan Xiong & Allison Koenecke & Michael Powell & Zhu Shen & Joshua T. Vogelstein & Susan Athey, 2021. "Federated Causal Inference in Heterogeneous Observational Data," Papers 2107.11732, arXiv.org, revised Apr 2023.
    2. Harsh Parikh & Trang Quynh Nguyen & Elizabeth A. Stuart & Kara E. Rudolph & Caleb H. Miles, 2025. "A Cautionary Tale on Integrating Studies with Disparate Outcome Measures for Causal Inference," Papers 2505.11014, arXiv.org.
    3. Ezinne Nwankwo & Lauri Goldkind & Angela Zhou, 2025. "Optimal Causal Annotations: An Application to Casenotes in Social Services," Papers 2502.10605, arXiv.org, revised Jul 2026.
    4. Jiafeng Chen & David M. Ritzwoller, 2021. "Semiparametric Estimation of Long-Term Treatment Effects," Papers 2107.14405, arXiv.org, revised Aug 2023.
    5. Aysegül Kayaoglu & Ghassan Baliki & Tilman Brück & Melodie Al Daccache & Dorothee Weiffen, 2023. "How to conduct impact evaluations in humanitarian and conflict settings," HiCN Working Papers 387, Households in Conflict Network.
    6. Dmitry Arkhangelsky & Guido Imbens, 2023. "Causal Models for Longitudinal and Panel Data: A Survey," Papers 2311.15458, arXiv.org, revised Jun 2024.
    7. Brett R. Gordon & Robert Moakler & Florian Zettelmeyer, 2023. "Predicted Incrementality by Experimentation (PIE) for Ad Measurement," Papers 2304.06828, arXiv.org, revised Apr 2026.
    8. Bokelmann, Björn & Lessmann, Stefan, 2025. "Heteroscedasticity-aware stratified sampling to improve uplift modeling," European Journal of Operational Research, Elsevier, vol. 325(1), pages 118-131.
    9. Harsh Parikh & Marco Morucci & Vittorio Orlandi & Sudeepa Roy & Cynthia Rudin & Alexander Volfovsky, 2023. "A Double Machine Learning Approach to Combining Experimental and Observational Data," Papers 2307.01449, arXiv.org, revised Oct 2025.
    10. Yacoubou Djima, Ismael & Kilic, Talip, 2024. "Attenuating measurement errors in agricultural productivity analysis by combining objective and self-reported survey data," Journal of Development Economics, Elsevier, vol. 168(C).
    11. Guido Imbens & Nathan Kallus & Xiaojie Mao & Yuhao Wang, 2022. "Long-term Causal Inference Under Persistent Confounding via Data Combination," Papers 2202.07234, arXiv.org, revised Aug 2024.
    12. Li, Ting & Shi, Chengchun & Wen, Qianglin & Sui, Yang & Qin, Yongli & Lai, Chunbo & Zhu, Hongtu, 2024. "Combining experimental and historical data for policy evaluation," LSE Research Online Documents on Economics 125588, London School of Economics and Political Science, LSE Library.
    13. Zhexiao Lin & Pablo Crespo, 2024. "Variance reduction combining pre-experiment and in-experiment data," Papers 2410.09027, arXiv.org, revised Mar 2026.
    14. Yechan Park & Yuya Sasaki, 2024. "A Bracketing Relationship for Long-Term Policy Evaluation with Combined Experimental and Observational Data," Papers 2401.12050, arXiv.org.
    15. Park, Yechan & Sasaki, Yuya, 2026. "The informativeness of combined experimental and observational data under dynamic selection," Journal of Econometrics, Elsevier, vol. 254(PB).
    16. Leah A. Jacobs & Alec McClean & Zach Branson & Edward Kennedy & Alex Fixler, 2024. "Incremental Propensity Score Effects for Criminology: An Application Assessing the Relationship Between Homelessness, Behavioral Health Problems, and Recidivism," Journal of Quantitative Criminology, Springer, vol. 40(4), pages 707-726, December.
    17. Mao, Minghai & Raiola, Antonio & Yang, Da, 2025. "Double machine learning for Oaxaca-Blinder decomposition," Economics Letters, Elsevier, vol. 255(C).
    18. Asanov, Anastasiya-Mariya & Asanov, Igor & Buenstorf, Guido, 2024. "A low-cost digital first aid tool to reduce psychological distress in refugees: A multi-country randomized controlled trial of self-help online in the first months after the invasion of Ukraine," Social Science & Medicine, Elsevier, vol. 362(C).
    19. Justin Whitehouse & Qizhao Chen & Morgane Austern & Vasilis Syrgkanis, 2025. "Inference on Optimal Policy Values and Other Irregular Functionals via Softmax Smoothing," Papers 2507.11780, arXiv.org, revised Mar 2026.
    20. Nicolaj N. Mühlbach, 2020. "Tree-based Synthetic Control Methods: Consequences of moving the US Embassy," CREATES Research Papers 2020-04, Department of Economics and Business Economics, Aarhus University.

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2606.17165. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.