IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2606.29784.html

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

Author

Listed:
  • Xinrui Ruan
  • Zhenyu Zhao
  • Waverly Wei
  • Yueshan Zhang
  • Zeyu Zheng
  • Sui Huang
  • Jingshen Wang

Abstract

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets.

Suggested Citation

  • Xinrui Ruan & Zhenyu Zhao & Waverly Wei & Yueshan Zhang & Zeyu Zheng & Sui Huang & Jingshen Wang, 2026. "HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data," Papers 2606.29784, arXiv.org.
  • Handle: RePEc:arx:papers:2606.29784
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2606.29784
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. A. P. Dawid & A. M. Skene, 1979. "Maximum Likelihood Estimation of Observer Error‐Rates Using the EM Algorithm," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 28(1), pages 20-28, March.
    2. Vaart,A. W. van der, 2000. "Asymptotic Statistics," Cambridge Books, Cambridge University Press, number 9780521784504.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Laurent Davezies & Xavier D'Haultfoeuille & Yannick Guyonvarch, 2019. "Empirical Process Results for Exchangeable Arrays," Papers 1906.11293, arXiv.org, revised May 2020.
    2. Shuang Liu, 2025. "Asymptotic Analysis of the Bias–Variance Trade-Off in Subsampling Metropolis–Hastings," Mathematics, MDPI, vol. 13(21), pages 1-30, October.
    3. Alexander Frankel & Maximilian Kasy, 2022. "Which Findings Should Be Published?," American Economic Journal: Microeconomics, American Economic Association, vol. 14(1), pages 1-38, February.
    4. Kasy, Maximilian, 2011. "A nonparametric test for path dependence in discrete panel data," Economics Letters, Elsevier, vol. 113(2), pages 172-175.
    5. Luofeng Liao & Christian Kroer, 2024. "Statistical Inference and A/B Testing in Fisher Markets and Paced Auctions," Papers 2406.15522, arXiv.org, revised Mar 2025.
    6. Waverly Wei & Maya Petersen & Mark J van der Laan & Zeyu Zheng & Chong Wu & Jingshen Wang, 2023. "Efficient targeted learning of heterogeneous treatment effects for multiple subgroups," Biometrics, The International Biometric Society, vol. 79(3), pages 1934-1946, September.
    7. Yao, Haixiang & Huang, Jinbo & Li, Yong & Humphrey, Jacquelyn E., 2021. "A general approach to smooth and convex portfolio optimization using lower partial moments," Journal of Banking & Finance, Elsevier, vol. 129(C).
    8. Jun‐ya Gotoh & Michael Jong Kim & Andrew E. B. Lim, 2021. "Calibration of Distributionally Robust Empirical Optimization Models," Operations Research, INFORMS, vol. 69(5), pages 1630-1650, September.
    9. Du, Mingyue & Zeng, Ricong, 2026. "Estimation of semiparametric probit model based on case-cohort interval-censored failure time data," Computational Statistics & Data Analysis, Elsevier, vol. 213(C).
    10. Yeganeh Alimohammadi & Karissa Huang & Christian Borgs & Jennifer Chayes, 2026. "Auditing the Auditors: Does Community-based Moderation Get It Right?," Papers 2603.18053, arXiv.org, revised May 2026.
    11. Jan-Lukas Wermuth, 2025. "Proper Correlation Coefficients for Nominal Random Variables," LIS Working papers 897, LIS Cross-National Data Center in Luxembourg.
    12. Luo, Yu & Graham, Daniel J. & McCoy, Emma J., 2023. "Semiparametric Bayesian doubly robust causal estimation," LSE Research Online Documents on Economics 117944, London School of Economics and Political Science, LSE Library.
    13. Ayden Higgins & Koen Jochmans, 2024. "Bootstrap Inference for Fixed‐Effect Models," Econometrica, Econometric Society, vol. 92(2), pages 411-427, March.
    14. Ashesh Rambachan & Jonathan Roth, 2026. "Design-Based Uncertainty for Quasi-Experiments," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 121(553), pages 477-491, January.
    15. Li, J. & Nott, D.J. & Fan, Y. & Sisson, S.A., 2017. "Extending approximate Bayesian computation methods to high dimensions via a Gaussian copula model," Computational Statistics & Data Analysis, Elsevier, vol. 106(C), pages 77-89.
    16. Higgins, Ayden & Jochmans, Koen, 2025. "Inference in Dynamic Models for Panel Data Using The Moving Block Bootstrap," TSE Working Papers 25-1620, Toulouse School of Economics (TSE).
    17. Denis Koshelev & Alexey Ponomarenko & Sergei Seleznev, 2023. "Amortized neural networks for agent-based model forecasting," Papers 2308.05753, arXiv.org.
    18. Debashis Ghosh, 2004. "Semiparametric methods for the binormal model with multiple biomarkers," The University of Michigan Department of Biostatistics Working Paper Series 1046, Berkeley Electronic Press.
    19. Yao Luo & Peijun Sang, 2022. "Efficient Estimation of Structural Models via Sieves," Papers 2204.13488, arXiv.org, revised Feb 2025.
    20. Brian D. Williamson & Peter B. Gilbert & Marco Carone & Noah Simon, 2021. "Nonparametric variable importance assessment using machine learning techniques," Biometrics, The International Biometric Society, vol. 77(1), pages 9-22, March.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2606.29784. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.