IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2606.11798.html

Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems

Author

Listed:
  • Xin Guo
  • Yijie Huang
  • Xiang Yu

Abstract

In this paper, we develop a continuous-time model-free reinforcement learning algorithm to learn deterministic equilibrium policies in general time-inconsistent control problems. Utilizing the extended Hamilton-Jacobi-Bellman system, we recast the original time-inconsistent problem into an equivalent two-stage problem. In the first stage, for given auxiliary functions, we employ the deterministic policy gradient approach to learn an optimal policy in an auxiliary time-consistent control problem. In the second stage, given the updated policy, we exploit the inner fixed point iterations and some martingale characterizations to learn the auxiliary functions. As a theoretical contribution, we provide some mild model assumptions and establish the convergence of inner fixed point iterations. By repeating this actor-critic style of iterations across two stages, our algorithm aims to learn the equilibrium under different sources of time-inconsistency in a unified manner. The superior effectiveness of the proposed algorithm are illustrated in two classical financial applications with time-inconsistency: mean-variance portfolio management and optimal tracking portfolio under non-exponential discounting.

Suggested Citation

  • Xin Guo & Yijie Huang & Xiang Yu, 2026. "Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems," Papers 2606.11798, arXiv.org.
  • Handle: RePEc:arx:papers:2606.11798
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2606.11798
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. R. H. Strotz, 1955. "Myopia and Inconsistency in Dynamic Utility Maximization," The Review of Economic Studies, Review of Economic Studies Ltd, vol. 23(3), pages 165-180.
    2. Min Dai & Yuchao Dong & Yanwei Jia, 2023. "Learning equilibrium mean‐variance strategy," Mathematical Finance, Wiley Blackwell, vol. 33(4), pages 1166-1212, October.
    3. Tomas Björk & Mariana Khapko & Agatha Murgoci, 2017. "On time-inconsistent stochastic control in continuous time," Finance and Stochastics, Springer, vol. 21(2), pages 331-360, April.
    4. Tomas Björk & Agatha Murgoci, 2014. "A theory of Markovian time-inconsistent stochastic control in discrete time," Finance and Stochastics, Springer, vol. 18(3), pages 545-592, July.
    5. Yanwei Jia & Xun Yu Zhou, 2021. "Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms," Papers 2111.11232, arXiv.org, revised Jul 2022.
    6. Tomas Björk & Mariana Khapko & Agatha Murgoci, 2021. "Time-Inconsistent Control Theory with Finance Applications," Springer Finance, Springer, number 978-3-030-81843-2, March.
    7. Yanwei Jia & Xun Yu Zhou, 2021. "Policy Evaluation and Temporal-Difference Learning in Continuous Time and Space: A Martingale Approach," Papers 2108.06655, arXiv.org, revised Feb 2022.
    8. Boyu Wang & Xuefeng Gao & Lingfei Li, 2026. "Reinforcement learning for continuous-time optimal execution: actor–critic algorithm and error analysis," Finance and Stochastics, Springer, vol. 30(2), pages 597-655, April.
    9. Xin Guo & Renyuan Xu & Thaleia Zariphopoulou, 2022. "Entropy Regularization for Mean Field Games with Learning," Mathematics of Operations Research, INFORMS, vol. 47(4), pages 3239-3260, November.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Yuchao Dong & Harry Zheng, 2025. "Extended HJB Equation for Mean-Variance Stopping Problem: Vanishing Regularization Method," Papers 2510.24128, arXiv.org.
    2. Bender, Christian & Thuan, Nguyen Tran, 2026. "Continuous time reinforcement learning: A random measure approach," Stochastic Processes and their Applications, Elsevier, vol. 194(C).
    3. Dylan Possamai & Mateo Rodriguez Polo, 2026. "Here, there and everywhere: state-dependent time-inconsistent stochastic control," Papers 2603.22022, arXiv.org.
    4. Fahrenwaldt, Matthias Albrecht & Jensen, Ninna Reitzel & Steffensen, Mogens, 2020. "Nonrecursive separation of risk and time preferences," Journal of Mathematical Economics, Elsevier, vol. 90(C), pages 95-108.
    5. Marcel Nutz & Yuchong Zhang, 2019. "Conditional Optimal Stopping: A Time-Inconsistent Optimization," Papers 1901.05802, arXiv.org, revised Oct 2019.
    6. Zhiping Chen & Liyuan Wang & Ping Chen & Haixiang Yao, 2019. "Continuous-Time Mean–Variance Optimization For Defined Contribution Pension Funds With Regime-Switching," International Journal of Theoretical and Applied Finance (IJTAF), World Scientific Publishing Co. Pte. Ltd., vol. 22(06), pages 1-33, September.
    7. Kang, Jian-hao & Gou, Zhun & Huang, Nan-jing, 2026. "Equilibrium reinsurance and investment strategies for insurers with random risk aversion under Heston’s SV model," Mathematics and Computers in Simulation (MATCOM), Elsevier, vol. 242(C), pages 343-365.
    8. Qian Lei & Chi Seng Pun, 2021. "Nonlocality, Nonlinearity, and Time Inconsistency in Stochastic Differential Games," Papers 2112.14409, arXiv.org, revised Sep 2023.
    9. Wang, Ling & Jia, Bowen, 2025. "Equilibrium investment strategies for a defined contribution pension plan with random risk aversion," Insurance: Mathematics and Economics, Elsevier, vol. 125(C).
    10. Soren Christensen & Kristoffer Lindensjo, 2019. "Moment constrained optimal dividends: precommitment \& consistent planning," Papers 1909.10749, arXiv.org.
    11. Soren Christensen & Kristoffer Lindensjo, 2019. "Time-inconsistent stopping, myopic adjustment & equilibrium stability: with a mean-variance application," Papers 1909.11921, arXiv.org, revised Jan 2020.
    12. Denis Belomestny & Tobias Hübner & Volker Krätschmer, 2022. "Solving optimal stopping problems under model uncertainty via empirical dual optimisation," Finance and Stochastics, Springer, vol. 26(3), pages 461-503, July.
    13. Bingyan Han & Chi Seng Pun & Hoi Ying Wong, 2021. "Robust state-dependent mean–variance portfolio selection: a closed-loop approach," Finance and Stochastics, Springer, vol. 25(3), pages 529-561, July.
    14. Erhan Bayraktar & Jingjie Zhang & Zhou Zhou, 2021. "Equilibrium concepts for time‐inconsistent stopping problems in continuous time," Mathematical Finance, Wiley Blackwell, vol. 31(1), pages 508-530, January.
    15. Xue Dong He & Xun Yu Zhou, 2021. "Who Are I: Time Inconsistency and Intrapersonal Conflict and Reconciliation," Papers 2105.01829, arXiv.org.
    16. Qian Lei & Chi Seng Pun, 2024. "A Malliavin Calculus Approach to Backward Stochastic Volterra Integral Equations," Papers 2412.19236, arXiv.org, revised Jan 2025.
    17. Bian, Lihua & Li, Zhongfei & Yao, Haixiang, 2018. "Pre-commitment and equilibrium investment strategies for the DC pension plan with regime switching and a return of premiums clause," Insurance: Mathematics and Economics, Elsevier, vol. 81(C), pages 78-94.
    18. Christensen, Sören & Lindensjö, Kristoffer, 2020. "On time-inconsistent stopping problems and mixed strategy stopping times," Stochastic Processes and their Applications, Elsevier, vol. 130(5), pages 2886-2917.
    19. Yuling Max Chen & Bin Li & David Saunders, 2025. "Exploratory Mean-Variance with Jumps: An Equilibrium Approach," Papers 2512.09224, arXiv.org.
    20. Samuel N. Cohen & Tanut Treetanthiploet, 2019. "Gittins' theorem under uncertainty," Papers 1907.05689, arXiv.org, revised Jun 2021.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2606.11798. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.