Divide-and-conquer offline policy evaluation for contextual bandits

Divide-and-conquer offline policy evaluation for contextual bandits

Author

Listed:

Wang, Weiwei
Shapovalova, Yuliya
Li, Yuqiang
Wu, Xianyi

Abstract

This paper investigates the application of divide-and-conquer (DC) algorithm to address the challenge of processing large datasets in offline policy evaluation within contextual bandit settings. We address the critical issue of determining the optimal number of machines as the dataset size scales, and establish a theoretical upper bound on the number of machines to control information loss from the DC algorithm. Our work aims at developing an estimator whose estimation accuracy matches that of an ideal direct estimator obtained by using the complete dataset. It turns out that the DC estimator can improve computational efficiency while maintaining statistical efficiency. When the number of machines is appropriately chosen, the estimator can be optimal in minimax rate. Furthermore, we extend the application of the DC algorithm to offline policy evaluation in reinforcement learning (RL) and explore the relationships between the number of machines and combinations of distribution shifts and horizons, showcasing enhanced computational efficiency through an extensive set of simulation experiments.

Suggested Citation

Wang, Weiwei & Shapovalova, Yuliya & Li, Yuqiang & Wu, Xianyi, 2025. "Divide-and-conquer offline policy evaluation for contextual bandits," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 676(C).

Handle: RePEc:eee:phsmap:v:676:y:2025:i:c:s0378437125004741
DOI: 10.1016/j.physa.2025.130822

Download full text from publisher

As the access to this document is restricted, you may want to

for a different version of it.

References listed on IDEAS

Zhan, Ruohan & Hadad, Vitor & Hirshberg, David A. & Athey, Susan, 2021. "Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits," Research Papers 3970, Stanford University, Graduate School of Business.
- Ruohan Zhan & Vitor Hadad & David A. Hirshberg & Susan Athey, 2021. "Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits," Papers 2106.02029, arXiv.org, revised Jun 2021.
Xi Chen & Jason D. Lee & He Li & Yun Yang, 2022. "Distributed Estimation for Principal Component Analysis: An Enlarged Eigenspace Analysis," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 117(540), pages 1775-1786, October.
Shie Mannor & Duncan Simester & Peng Sun & John N. Tsitsiklis, 2007. "Bias and Variance Approximation in Value Function Estimates," Management Science, INFORMS, vol. 53(2), pages 308-322, February.
Doudou Zhou & Yufeng Zhang & Aaron Sonabend-W & Zhaoran Wang & Junwei Lu & Tianxi Cai, 2024. "Federated Offline Reinforcement Learning," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 119(548), pages 3152-3163, October.
Xin Zhou & Nicole Mayer-Hamblett & Umer Khan & Michael R. Kosorok, 2017. "Residual Weighted Learning for Estimating Individualized Treatment Rules," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 112(517), pages 169-187, January.

Full references (including those not matched with items on IDEAS)

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Ruohan Zhan & Zhimei Ren & Susan Athey & Zhengyuan Zhou, 2024. "Policy Learning with Adaptively Collected Data," Management Science, INFORMS, vol. 70(8), pages 5270-5297, August.
- Zhan, Ruohan & Ren, Zhimei & Athey, Susan & Zhou, Zhengyuan, 2021. "Policy Learning with Adaptively Collected Data," Research Papers 3963, Stanford University, Graduate School of Business.
- Ruohan Zhan & Zhimei Ren & Susan Athey & Zhengyuan Zhou, 2021. "Policy Learning with Adaptively Collected Data," Papers 2105.02344, arXiv.org, revised Nov 2022.
Q. Clairon & R. Henderson & N. J. Young & E. D. Wilson & C. J. Taylor, 2021. "Adaptive treatment and robust control," Biometrics, The International Biometric Society, vol. 77(1), pages 223-236, March.
Pedro Afonso Fernandes, 2024. "Forecasting with Neuro-Dynamic Programming," Papers 2404.03737, arXiv.org.
Zhengyuan Zhou & Susan Athey & Stefan Wager, 2023. "Offline Multi-Action Policy Learning: Generalization and Optimization," Operations Research, INFORMS, vol. 71(1), pages 148-183, January.
- Zhou, Zhengyuan & Athey, Susan & Wager, Stefan, 2018. "Offline Multi-Action Policy Learning: Generalization and Optimization," Research Papers 3734, Stanford University, Graduate School of Business.
- Zhengyuan Zhou & Susan Athey & Stefan Wager, 2018. "Offline Multi-Action Policy Learning: Generalization and Optimization," Papers 1810.04778, arXiv.org, revised Nov 2018.
Crystal T. Nguyen & Daniel J. Luckett & Anna R. Kahkoska & Grace E. Shearrer & Donna Spruijt‐Metz & Jaimie N. Davis & Michael R. Kosorok, 2020. "Estimating individualized treatment regimes from crossover designs," Biometrics, The International Biometric Society, vol. 76(3), pages 778-788, September.
Jinglong Zhao, 2024. "Experimental Design For Causal Inference Through An Optimization Lens," Papers 2408.09607, arXiv.org, revised Aug 2024.
Kushal S. Shah & Haoda Fu & Michael R. Kosorok, 2023. "Stabilized direct learning for efficient estimation of individualized treatment rules," Biometrics, The International Biometric Society, vol. 79(4), pages 2843-2856, December.
Jonas Metzger, 2022. "Adversarial Estimators," Papers 2204.10495, arXiv.org, revised Jun 2022.
Shiau Hong Lim & Huan Xu & Shie Mannor, 2016. "Reinforcement Learning in Robust Markov Decision Processes," Mathematics of Operations Research, INFORMS, vol. 41(4), pages 1325-1353, November.
Shuze Chen & David Simchi-Levi & Chonghuan Wang, 2024. "Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality," Papers 2407.19618, arXiv.org, revised Sep 2025.
Vasilis Syrgkanis & Ruohan Zhan, 2023. "Post Reinforcement Learning Inference," Papers 2302.08854, arXiv.org, revised Oct 2025.
Brian Cho & Ana-Roxana Pop & Ariel Evnine & Nathan Kallus, 2025. "SNPL: Simultaneous Policy Learning and Evaluation for Safe Multi-Objective Policy Improvement," Papers 2503.12760, arXiv.org, revised Mar 2025.
Susan Athey & Stefan Wager, 2021. "Policy Learning With Observational Data," Econometrica, Econometric Society, vol. 89(1), pages 133-161, January.
- Susan Athey & Stefan Wager, 2017. "Policy Learning with Observational Data," Papers 1702.02896, arXiv.org, revised Sep 2020.
Weibin Mo & Yufeng Liu, 2022. "Efficient learning of optimal individualized treatment rules for heteroscedastic or misspecified treatment‐free effect models," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 84(2), pages 440-472, April.
Wolfram Wiesemann & Daniel Kuhn & Berç Rustem, 2010. "Robust Markov Decision Processes," Working Papers 034, COMISEF.
Varagapriya, V & Singh, Vikas Vikram & Lisser, Abdel, 2024. "Rank-1 transition uncertainties in constrained Markov decision processes," European Journal of Operational Research, Elsevier, vol. 318(1), pages 167-178.
Shie Mannor & Ofir Mebel & Huan Xu, 2016. "Robust MDPs with k -Rectangular Uncertainty," Mathematics of Operations Research, INFORMS, vol. 41(4), pages 1484-1509, November.
Yi Zhu & Jing Dong & Henry Lam, 2024. "Uncertainty Quantification and Exploration for Reinforcement Learning," Operations Research, INFORMS, vol. 72(4), pages 1689-1709, July.
Omid Rafieian, 2023. "Optimizing User Engagement Through Adaptive Ad Sequencing," Marketing Science, INFORMS, vol. 42(5), pages 910-933, September.
Ilbin Lee, 2024. "Is Separately Modeling Subpopulations Beneficial for Sequential Decision-Making?," Operations Research, INFORMS, vol. 72(6), pages 2595-2611, November.

More about this item

Keywords

; ; ; ; ;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:phsmap:v:676:y:2025:i:c:s0378437125004741. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.journals.elsevier.com/physica-a-statistical-mechpplications/ .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Divide-and-conquer offline policy evaluation for contextual bandits

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data