IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2606.17545.html

Continuous-time Optimal Stopping through Deep Reinforcement Learning

Author

Listed:
  • Cosmin Borsa
  • Michael Ludkovski

Abstract

Simulation based solvers for optimal stopping problems must discretize the stopping decision. Under classical dynamic programming, a coarse exercise grid with only a few stopping opportunities can materially undervalue the optimal expected reward, whereas on a very fine grid, approximation errors accumulate through the backward recursion. To remove this limitation, we develop a new reinforcement-learning inspired algorithm that enables us to learn the exercise rule at arbitrarily fine time resolution. Our CARLOS (Continuous-time Adaptive Reinforcement Learning for Optimal Stopping) algorithm utilizes an aggregate deep neural network (ADNN) to learn a joint space-time decision boundary. Starting from a coarse time grid, we progressively increase the frequency of stopping opportunities, while in parallel training the ADNN to refine its timing-value estimates. We moreover design an adaptive sampling strategy that gradually concentrates training effort near the stopping boundary. Benchmarked results show that CARLOS delivers higher prices than existing Bermudan solvers, approaching the American upper bound, and achieves high computational efficiency relative to non-RL comparators.

Suggested Citation

  • Cosmin Borsa & Michael Ludkovski, 2026. "Continuous-time Optimal Stopping through Deep Reinforcement Learning," Papers 2606.17545, arXiv.org.
  • Handle: RePEc:arx:papers:2606.17545
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2606.17545
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Leif Andersen & Mark Broadie, 2004. "Primal-Dual Simulation Algorithm for Pricing Multidimensional American Options," Management Science, INFORMS, vol. 50(9), pages 1222-1234, September.
    2. Longstaff, Francis A & Schwartz, Eduardo S, 2001. "Valuing American Options by Simulation: A Simple Least-Squares Approach," The Review of Financial Studies, Society for Financial Studies, vol. 14(1), pages 113-147.
    3. Bernard Lapeyre & Jérôme Lelong, 2021. "Neural network regression for Bermudan option pricing," Post-Print hal-02183587, HAL.
    4. Rongju Zhang & Nicolas Langrené & Yu Tian & Zili Zhu & Fima Klebaner & Kais Hamza, 2019. "Dynamic portfolio optimization with liquidity cost and market impact: a simulation-and-regression approach," Quantitative Finance, Taylor & Francis Journals, vol. 19(3), pages 519-532, March.
    5. Mike Ludkovski, 2022. "Regression Monte Carlo for Impulse Control," Papers 2203.06539, arXiv.org.
    6. Sebastian Becker & Patrick Cheridito & Arnulf Jentzen & Timo Welti, 2019. "Solving high-dimensional optimal stopping problems using deep learning," Papers 1908.01602, arXiv.org, revised Aug 2021.
    7. Justin Sirignano & Konstantinos Spiliopoulos, 2017. "DGM: A deep learning algorithm for solving partial differential equations," Papers 1708.07469, arXiv.org, revised Sep 2018.
    8. Lukas Gonon, 2024. "Deep neural network expressivity for optimal stopping problems," Finance and Stochastics, Springer, vol. 28(3), pages 865-910, July.
    9. John Ery & Loris Michel, 2021. "Solving optimal stopping problems with Deep Q-Learning," Papers 2101.09682, arXiv.org, revised Jun 2024.
    10. Yangang Chen & Justin W. L. Wan, 2021. "Deep neural network framework based on backward stochastic differential equations for pricing and hedging American options in high dimensions," Quantitative Finance, Taylor & Francis Journals, vol. 21(1), pages 45-67, January.
    11. Jiefei Yang & Guanglian Li, 2024. "A deep primal-dual BSDE method for optimal stopping problems," Papers 2409.06937, arXiv.org.
    12. Rongju Zhang & Nicolas Langrené & Yu Tian & Zili Zhu & Fima Klebaner & Kais Hamza, 2019. "Dynamic portfolio optimization with liquidity cost and market impact: a simulation-and-regression approach," Post-Print hal-02909207, HAL.
    13. Ruimeng Hu, 2020. "Deep learning for ranking response surfaces with applications to optimal stopping problems," Quantitative Finance, Taylor & Francis Journals, vol. 20(9), pages 1567-1581, September.
    14. Ivan Guo & Nicolas Langrené & Jiahao Wu, 2025. "Simultaneous upper and lower bounds of American-style option prices with hedging via neural networks," Quantitative Finance, Taylor & Francis Journals, vol. 25(4), pages 509-525, April.
    15. Longstaff, Francis A & Schwartz, Eduardo S, 2001. "Valuing American Options by Simulation: A Simple Least-Squares Approach," University of California at Los Angeles, Anderson Graduate School of Management qt43n1k4jb, Anderson Graduate School of Management, UCLA.
    16. Andres Max Reppen & Halil Mete Soner & Valentin Tissot‐Daguette, 2025. "Neural optimal stopping boundary," Mathematical Finance, Wiley Blackwell, vol. 35(2), pages 441-469, April.
    17. Rongju Zhang & Nicolas Langren'e & Yu Tian & Zili Zhu & Fima Klebaner & Kais Hamza, 2016. "Dynamic portfolio optimization with liquidity cost and market impact: a simulation-and-regression approach," Papers 1610.07694, arXiv.org, revised Jun 2019.
    18. Ruimeng Hu, 2019. "Deep Learning for Ranking Response Surfaces with Applications to Optimal Stopping Problems," Papers 1901.03478, arXiv.org, revised Mar 2020.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Daniel Chee & Noufel Frikha & Libo Li, 2026. "Entropy-regularized penalization schemes for American options and reflected BSDEs with singular generators," Université Paris1 Panthéon-Sorbonne (Post-Print and Working Papers) hal-05520660, HAL.
    2. Lukas Gonon, 2022. "Deep neural network expressivity for optimal stopping problems," Papers 2210.10443, arXiv.org.
    3. Noufel Frikha & Libo Li & Daniel Chee, 2025. "An Entropy Regularized BSDE Approach to Bermudan Options and Games," Université Paris1 Panthéon-Sorbonne (Post-Print and Working Papers) hal-05265653, HAL.
    4. A. Max Reppen & H. Mete Soner & Valentin Tissot-Daguette, 2022. "Deep Stochastic Optimization in Finance," Papers 2205.04604, arXiv.org.
    5. Xuwei Yang & Anastasis Kratsios & Florian Krach & Matheus Grasselli & Aurelien Lucchi, 2023. "Regret-Optimal Federated Transfer Learning for Kernel Regression with Applications in American Option Pricing," Papers 2309.04557, arXiv.org, revised Oct 2024.
    6. Daniel Chee & Noufel Frikha & Libo Li, 2026. "A Monotone Limit Approach to Entropy-Regularized American Options," Papers 2602.18062, arXiv.org.
    7. Lukas Gonon, 2024. "Deep neural network expressivity for optimal stopping problems," Finance and Stochastics, Springer, vol. 28(3), pages 865-910, July.
    8. A. Max Reppen & H. Mete Soner & Valentin Tissot-Daguette, 2023. "Deep stochastic optimization in finance," Digital Finance, Springer, vol. 5(1), pages 91-111, March.
    9. repec:tin:wpaper:20230016 is not listed on IDEAS
    10. Rongju Zhang & Nicolas Langrené & Yu Tian & Zili Zhu & Fima Klebaner & Kais Hamza, 2019. "Skewed target range strategy for multiperiod portfolio optimization using a two-stage least squares Monte Carlo method," Post-Print hal-02909342, HAL.
    11. A. Max Reppen & H. Mete Soner & Valentin Tissot-Daguette, 2022. "Neural Optimal Stopping Boundary," Papers 2205.04595, arXiv.org, revised May 2023.
    12. Xiaoqiang Cai & Gen Yu, 2025. "Bayesian learning in dynamic portfolio selection under a minimax rule," OR Spectrum: Quantitative Approaches in Management, Springer;Gesellschaft für Operations Research e.V., vol. 47(1), pages 287-324, March.
    13. Daniel Chee & Noufel Frikha & Libo Li, 2026. "A Monotone Limit Approach to Entropy-Regularized American Options," Université Paris1 Panthéon-Sorbonne (Post-Print and Working Papers) hal-05520656, HAL.
    14. Jiefei Yang & Guanglian Li, 2024. "A deep primal-dual BSDE method for optimal stopping problems," Papers 2409.06937, arXiv.org.
    15. Ivan Guo & Nicolas Langrené & Gregoire Loeper & Wei Ning, 2020. "Robust utility maximization under model uncertainty via a penalization approach," Working Papers hal-02910261, HAL.
    16. Zineb El Filali Ech-Chafiq & Pierre Henry-Labordere & Jérôme Lelong, 2021. "Pricing Bermudan options using regression trees/random forests," Working Papers hal-03436046, HAL.
    17. Sebastian Becker & Patrick Cheridito & Arnulf Jentzen, 2020. "Pricing and Hedging American-Style Options with Deep Learning," JRFM, MDPI, vol. 13(7), pages 1-12, July.
    18. Ivan Guo & Nicolas Langren'e & Jiahao Wu, 2023. "Simultaneous upper and lower bounds of American-style option prices with hedging via neural networks," Papers 2302.12439, arXiv.org, revised Nov 2024.
    19. Chinonso Nwankwo & Nneka Umeorah & Tony Ware & Weizhong Dai, 2022. "Deep learning and American options via free boundary framework," Papers 2211.11803, arXiv.org, revised Dec 2022.
    20. Jasper Rou, 2025. "Time Deep Gradient Flow Method for pricing American options," Papers 2507.17606, arXiv.org.
    21. Vikranth Lokeshwar Dhandapani & Shashi Jain, 2024. "Optimizing Neural Networks for Bermudan Option Pricing: Convergence Acceleration, Future Exposure Evaluation and Interpolation in Counterparty Credit Risk," Papers 2402.15936, arXiv.org.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2606.17545. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.