IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2607.09461.html

Deep Learning for Dynamic Programming with Recursive Utility Using First-order Conditions

Author

Listed:
  • Xianhua Peng
  • Wu Guo
  • Songyan Wang
  • Jianfei Zhu

Abstract

This paper proposes the certainty-equivalent first-order learning (CEFOL) algorithm, a deep learning algorithm for solving discrete-time dynamic programming problems with recursive utility. Dynamic programming with recursive utility is challenging because nonlinear certainty equivalent appears in the Bellman equation and the first-order optimality conditions but is difficult to evaluate. By introducing a separate neural network to represent the certainty equivalent, CEFOL enables the exploitation of the Bellman and model-specific first-order optimality conditions. In addition to certainty equivalent, CEFOL also uses neural networks to learn the value functions, policy functions, and Lagrange multipliers by using model-specific first-order conditions to construct residuals for minimization. By using first-order and KKT residuals to learn the policy, CEFOL directly accommodates general equality and inequality constraints on the controls, including occasionally binding constraints, without requiring penalty functions or problem-specific reformulations. We apply the algorithm to risk-sensitive and Epstein--Zin consumption-saving problems, a small-noise robust-control problem, and a DSGE model with recursive preferences and stochastic volatility. Across these applications, out-of-sample Bellman diagnostics and model-specific optimality residuals, including Euler or first-order residuals where applicable, are generally of order 1.0e-4 to 1.0e-3 over the relevant state regions, with larger values mainly near binding constraints, and the learned value and policy functions closely match VFI benchmarks when available. The CEFOL algorithm also works for dynamic programming problems with expected utility, as expected utility is a special case of recursive utility.

Suggested Citation

  • Xianhua Peng & Wu Guo & Songyan Wang & Jianfei Zhu, 2026. "Deep Learning for Dynamic Programming with Recursive Utility Using First-order Conditions," Papers 2607.09461, arXiv.org.
  • Handle: RePEc:arx:papers:2607.09461
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2607.09461
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Ma, Qingyin & Stachurski, John & Toda, Alexis Akira, 2022. "Unbounded dynamic programming via the Q-transform," Journal of Mathematical Economics, Elsevier, vol. 100(C).
    2. Yongyang Cai & Thomas S. Lontzek, 2019. "The Social Cost of Carbon with Economic and Climate Risks," Journal of Political Economy, University of Chicago Press, vol. 127(6), pages 2684-2734.
    3. Dumas, Bernard & Uppal, Raman & Wang, Tan, 2000. "Efficient Intertemporal Allocations with Recursive Utility," Journal of Economic Theory, Elsevier, vol. 93(2), pages 240-259, August.
    4. Duffie, Darrel & Lions, Pierre-Louis, 1992. "PDE solutions of stochastic differential utility," Journal of Mathematical Economics, Elsevier, vol. 21(6), pages 577-606.
    5. Marlon Azinovic & Jan v{Z}emliv{c}ka, 2023. "Economics-Inspired Neural Networks with Stabilizing Homotopies," Papers 2303.14802, arXiv.org.
    6. Vytautas Valaitis & Alessandro T. Villa, 2024. "A machine learning projection method for macro‐finance models," Quantitative Economics, Econometric Society, vol. 15(1), pages 145-173, January.
    7. John Y. Campbell & Yeung Lewis Chanb & M. Viceira, 2013. "A multivariate model of strategic asset allocation," World Scientific Book Chapters, in: Leonard C MacLean & William T Ziemba (ed.), HANDBOOK OF THE FUNDAMENTALS OF FINANCIAL DECISION MAKING Part II, chapter 39, pages 809-848, World Scientific Publishing Co. Pte. Ltd..
    8. Lars Peter Hansen & Thomas J. Sargent, 2013. "Recursive Models of Dynamic Linear Economies," Economics Books, Princeton University Press, edition 1, number 10141, December.
    9. Emmet Hall-Hoffarth, 2023. "Non-linear approximations of DSGE models with neural-networks and hard-constraints," Papers 2310.13436, arXiv.org.
    10. Weil, Philippe, 1989. "The equity premium puzzle and the risk-free rate puzzle," Journal of Monetary Economics, Elsevier, vol. 24(3), pages 401-421, November.
    11. TallariniJr., Thomas D., 2000. "Risk-sensitive real business cycles," Journal of Monetary Economics, Elsevier, vol. 45(3), pages 507-532, June.
    12. Victor Duarte & Diogo Duarte & Dejanir H Silva, 2024. "Machine Learning for Continuous-Time Finance," The Review of Financial Studies, Society for Financial Studies, vol. 37(11), pages 3217-3271.
    13. Victor Duarte & Julia Fonseca & Aaron S. Goodman & Jonathan A. Parker, 2021. "Simple Allocation Rules and Optimal Portfolio Choice Over the Lifecycle," NBER Working Papers 29559, National Bureau of Economic Research, Inc.
    14. Amine Mohamed Aboussalah & Ziyun Xu & Chi-Guhn Lee, 2022. "What is the value of the cross-sectional approach to deep reinforcement learning?," Quantitative Finance, Taylor & Francis Journals, vol. 22(6), pages 1091-1111, June.
    15. Walter Pohl & Karl Schmedders & Ole Wilms, 2024. "Existence of the Wealth-Consumption Ratio in Asset Pricing Models with Recursive Preferences," The Review of Financial Studies, Society for Financial Studies, vol. 37(3), pages 989-1028.
    16. Achref Bachouch & Côme Huré & Nicolas Langrené & Huyên Pham, 2022. "Deep Neural Networks Algorithms for Stochastic Control Problems on Finite Horizon: Numerical Applications," Methodology and Computing in Applied Probability, Springer, vol. 24(1), pages 143-178, March.
    17. Kenneth L. Judd, 1998. "Numerical Methods in Economics," MIT Press Books, The MIT Press, edition 1, volume 1, number 0262100711, December.
    18. Qingyin Ma & John Stachurski, 2021. "Dynamic Programming Deconstructed: Transformations of the Bellman Equation and Computational Efficiency," Operations Research, INFORMS, vol. 69(5), pages 1591-1607, September.
    19. Marlon Azinovic & Luca Gaegauf & Simon Scheidegger, 2022. "Deep Equilibrium Nets," International Economic Review, Department of Economics, University of Pennsylvania and Osaka University Institute of Social and Economic Research Association, vol. 63(4), pages 1471-1525, November.
    20. Anis Matoussi & Hao Xing, 2018. "Convex duality for Epstein–Zin stochastic differential utility," Mathematical Finance, Wiley Blackwell, vol. 28(4), pages 991-1019, October.
    21. Gaetano Bloise & Cuong Le Van & Yiannis Vailakis, 2024. "Do not Blame Bellman: It Is Koopmans' Fault," Econometrica, Econometric Society, vol. 92(1), pages 111-140, January.
    22. Christensen, Timothy M., 2022. "Existence and uniqueness of recursive utilities without boundedness," Journal of Economic Theory, Elsevier, vol. 200(C).
    23. Yifan Zhao & Arnab Basu & Thomas S. Lontzek & Karl Schmedders, 2023. "The Social Cost of Carbon When We Wish for Full-Path Robustness," Management Science, INFORMS, vol. 69(12), pages 7585-7606, December.
    24. Holger Kraft & Thomas Seiferling & Frank Thomas Seifried, 2017. "Optimal consumption and investment with Epstein–Zin recursive utility," Finance and Stochastics, Springer, vol. 21(1), pages 187-226, January.
    25. Duffie, Darrell & Epstein, Larry G, 1992. "Stochastic Differential Utility," Econometrica, Econometric Society, vol. 60(2), pages 353-394, March.
    26. Holger Kraft & Frank Seifried & Mogens Steffensen, 2013. "Consumption-portfolio optimization with recursive utility in incomplete markets," Finance and Stochastics, Springer, vol. 17(1), pages 161-196, January.
    27. Hao Xing, 2017. "Consumption–investment optimization with Epstein–Zin utility in incomplete markets," Finance and Stochastics, Springer, vol. 21(1), pages 227-262, January.
    28. Manuel S. Santos, 2000. "Accuracy of Numerical Solutions using the Euler Equation Residuals," Econometrica, Econometric Society, vol. 68(6), pages 1377-1402, November.
    29. Victor Duarte & Diogo Duarte & Dejanir H. Silva, 2024. "Machine Learning for Continuous-Time Finance," CESifo Working Paper Series 10909, CESifo.
    30. John Y. Campbell & Luis M. Viceira, 1999. "Consumption and Portfolio Decisions when Expected Returns are Time Varying," The Quarterly Journal of Economics, President and Fellows of Harvard College, vol. 114(2), pages 433-495.
    31. repec:bla:jfinan:v:59:y:2004:i:4:p:1481-1509 is not listed on IDEAS
    32. Philippe Weil, 1990. "Nonexpected Utility in Macroeconomics," The Quarterly Journal of Economics, President and Fellows of Harvard College, vol. 105(1), pages 29-42.
    33. Greg Kaplan & Giovanni L. Violante, 2014. "A Model of the Consumption Response to Fiscal Stimulus Payments," Econometrica, Econometric Society, vol. 82(4), pages 1199-1239, July.
    34. Judd, Kenneth L., 1992. "Projection methods for solving aggregate growth models," Journal of Economic Theory, Elsevier, vol. 58(2), pages 410-452, December.
    35. Kreps, David M & Porteus, Evan L, 1978. "Temporal Resolution of Uncertainty and Dynamic Choice Theory," Econometrica, Econometric Society, vol. 46(1), pages 185-200, January.
    36. Stachurski, John & Wilms, Ole & Zhang, Junnan, 2024. "Asset pricing with time preference shocks: Existence and uniqueness," Journal of Economic Theory, Elsevier, vol. 216(C).
    37. Gaetano Bloise & Cuong Le Van & Yiannis Vailakis, 2024. "Do not Blame Bellman: It Is Koopmans' Fault," PSE-Ecole d'économie de Paris (Postprint) halshs-04928565, HAL.
    38. Stachurski, John & Zhang, Junnan, 2021. "Dynamic programming with state-dependent discounting," Journal of Economic Theory, Elsevier, vol. 192(C).
    39. Martin Andreasen, 2012. "On the Effects of Rare Disasters and Uncertainty Shocks for Risk Premia in Non-Linear DSGE Models," Review of Economic Dynamics, Elsevier for the Society for Economic Dynamics, vol. 15(3), pages 295-316, July.
    40. Gaetano Bloise & Cuong Le Van & Yiannis Vailakis, 2024. "Do not Blame Bellman: It Is Koopmans' Fault," Post-Print halshs-04928565, HAL.
    41. Duffie, Darrell & Epstein, Larry G, 1992. "Asset Pricing with Stochastic Differential Utility," The Review of Financial Studies, Society for Financial Studies, vol. 5(3), pages 411-436.
    42. Xianhua Peng & Wu Guo, 2026. "Deep Learning for Dynamic Programming with Recursive Utility," Papers 2607.04278, arXiv.org.
    43. repec:spo:wpmain:info:hdl:2441/8686 is not listed on IDEAS
    44. Dario Caldara & Jesus Fernandez-Villaverde & Juan Rubio-Ramirez & Wen Yao, 2012. "Computing DSGE Models with Recursive Preferences and Stochastic Volatility," Review of Economic Dynamics, Elsevier for the Society for Economic Dynamics, vol. 15(2), pages 188-206, April.
    45. Stachurski, John & Wilms, Ole & Zhang, Junnan, 2024. "Asset pricing with time preference shocks: Existence and uniqueness," Other publications TiSEM 29da00af-3cca-4717-aa55-a, Tilburg University, School of Economics and Management.
    46. Schroder, Mark & Skiadas, Costis, 1999. "Optimal Consumption and Portfolio Selection with Stochastic Differential Utility," Journal of Economic Theory, Elsevier, vol. 89(1), pages 68-126, November.
    47. Maliar, Lilia & Maliar, Serguei & Winant, Pablo, 2021. "Deep learning for solving dynamic economic models," Journal of Monetary Economics, Elsevier, vol. 122(C), pages 76-101.
    48. Carroll, Christopher D., 2006. "The method of endogenous gridpoints for solving dynamic stochastic optimization problems," Economics Letters, Elsevier, vol. 91(3), pages 312-320, June.
    49. Matoussi, Anis & Xing, Hao, 2018. "Convex duality for Epstein-Zin stochastic differential utility," LSE Research Online Documents on Economics 82519, London School of Economics and Political Science, LSE Library.
    50. A. Max Reppen & H. Mete Soner & Valentin Tissot-Daguette, 2023. "Deep stochastic optimization in finance," Digital Finance, Springer, vol. 5(1), pages 91-111, March.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Xianhua Peng & Wu Guo, 2026. "Deep Learning for Dynamic Programming with Recursive Utility," Papers 2607.04278, arXiv.org.
    2. Dirk Becherer & Wilfried Kuissi-Kamdem & Olivier Menoukeu-Pamen, 2023. "Optimal consumption with labor income and borrowing constraints for recursive preferences," Working Papers hal-04017143, HAL.
    3. Dejian Tian & Weidong Tian & Jianjun Zhou & Zimu Zhu, 2025. "Optimal Consumption-Investment with Epstein-Zin Utility under Leverage Constraint," Papers 2509.21929, arXiv.org, revised Oct 2025.
    4. Li, Hanwu & Riedel, Frank & Yang, Shuzhen, 2024. "Optimal consumption for recursive preferences with local substitution — the case of certainty," Journal of Mathematical Economics, Elsevier, vol. 110(C).
    5. Xianhua Peng & Steven Kou & Lekang Zhang, 2024. "A Machine Learning Algorithm for Finite-Horizon Stochastic Control Problems in Economics," Papers 2411.08668, arXiv.org, revised Dec 2024.
    6. Zixin Feng & Dejian Tian, 2021. "Optimal consumption and portfolio selection with Epstein-Zin utility under general constraints," Papers 2111.09032, arXiv.org, revised May 2023.
    7. Joshua Aurand & Yu-Jui Huang, 2019. "Epstein-Zin Utility Maximization on a Random Horizon," Papers 1903.08782, arXiv.org, revised May 2023.
    8. Joshua Aurand & Yu‐Jui Huang, 2023. "Epstein‐Zin utility maximization on a random horizon," Mathematical Finance, Wiley Blackwell, vol. 33(4), pages 1370-1411, October.
    9. Feng, Zixin & Tian, Dejian & Zheng, Harry, 2026. "Consumption–investment optimization with Epstein–Zin utility in unbounded non-Markovian markets," Stochastic Processes and their Applications, Elsevier, vol. 192(C).
    10. Roche, Hervé, 2011. "Asset prices in an exchange economy when agents have heterogeneous homothetic recursive preferences and no risk free bond is available," Journal of Economic Dynamics and Control, Elsevier, vol. 35(1), pages 80-96, January.
    11. Campani, Carlos Heitor & Garcia, René, 2019. "Approximate analytical solutions for consumption/investment problems under recursive utility and finite horizon," The North American Journal of Economics and Finance, Elsevier, vol. 48(C), pages 364-384.
    12. Chen, Xingjiang & Ruan, Xinfeng & Zhang, Wenjun, 2021. "Dynamic portfolio choice and information trading with recursive utility," Economic Modelling, Elsevier, vol. 98(C), pages 154-167.
    13. Dariusz Zawisza, 2020. "On the parabolic equation for portfolio problems," Papers 2003.13317, arXiv.org, revised Oct 2020.
    14. Matoussi, Anis & Xing, Hao, 2018. "Convex duality for Epstein-Zin stochastic differential utility," LSE Research Online Documents on Economics 82519, London School of Economics and Political Science, LSE Library.
    15. Anis Matoussi & Hao Xing, 2016. "Convex duality for stochastic differential utility," Papers 1601.03562, arXiv.org.
    16. Guanxing Fu & Ulrich Horst, 2025. "Mean Field Portfolio Games with Epstein-Zin Preferences," Papers 2505.07231, arXiv.org, revised Jun 2026.
    17. Erhan Bayraktar & Emmet Lawless, 2026. "Infinite Horizon Optimal Consumption: Intertemporal Hedging under Epstein-Zin Preferences," Papers 2606.02945, arXiv.org, revised Aug 2026.
    18. Shigeta, Yuki, 2020. "Gain/loss asymmetric stochastic differential utility," Journal of Economic Dynamics and Control, Elsevier, vol. 118(C).
    19. Immacolata Oliva & Ilaria Stefani, 2023. "Co-jumps and recursive preferences in portfolio choices," Annals of Finance, Springer, vol. 19(3), pages 291-324, September.
    20. Kraft, Holger & Seifried, Frank Thomas, 2014. "Stochastic differential utility as the continuous-time limit of recursive utility," Journal of Economic Theory, Elsevier, vol. 151(C), pages 528-550.

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2607.09461. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.