IDEAS home Printed from https://ideas.repec.org/a/gam/jmathe/v13y2025i5p833-d1603779.html

Evolutionary Reinforcement Learning: A Systematic Review and Future Directions

Author

Listed:
  • Yuanguo Lin

    (School of Computer Engineering, Jimei University, Xiamen 361021, China)

  • Fan Lin

    (School of Informatics, Xiamen University, Xiamen 361005, China)

  • Guorong Cai

    (School of Computer Engineering, Jimei University, Xiamen 361021, China)

  • Hong Chen

    (Information Hub, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China)

  • Linxin Zou

    (School of Cyber Science and Engineering, Wuhan University, Wuhan 430074, China)

  • Yunxuan Liu

    (School of Computer Engineering, Jimei University, Xiamen 361021, China)

  • Pengcheng Wu

    (Webank-NTU Joint Research Institute on Fintech, Nanyang Technological University, Singapore 639798, Singapore)

Abstract

In response to the limitations of reinforcement learning and Evolutionary Algorithms (EAs) in complex problem-solving, Evolutionary Reinforcement Learning (EvoRL) has emerged as a synergistic solution. This systematic review aims to provide a comprehensive analysis of EvoRL, examining the symbiotic relationship between EAs and reinforcement learning algorithms and identifying critical gaps in relevant application tasks. The review begins by outlining the technological foundations of EvoRL, detailing the complementary relationship between EAs and reinforcement learning algorithms to address the limitations of reinforcement learning, such as parameter sensitivity, sparse rewards, and its susceptibility to local optima. We then delve into the challenges faced by both reinforcement learning and EvoRL, exploring the utility and limitations of EAs in EvoRL. EvoRL itself is constrained by the sampling efficiency and algorithmic complexity, which affect its application in areas like robotic control and large-scale industrial settings. Furthermore, we address significant open issues in the field, such as adversarial robustness, fairness, and ethical considerations. Finally, we propose future directions for EvoRL, emphasizing research avenues that strive to enhance self-adaptation, self-improvement, scalability, interpretability, and so on. To quantify the current state, we analyzed about 100 EvoRL studies, categorizing them based on algorithms, performance metrics, and benchmark tasks. Serving as a comprehensive resource for researchers and practitioners, this systematic review provides insights into the current state of EvoRL and offers a guide for advancing its capabilities in the ever-evolving landscape of artificial intelligence.

Suggested Citation

  • Yuanguo Lin & Fan Lin & Guorong Cai & Hong Chen & Linxin Zou & Yunxuan Liu & Pengcheng Wu, 2025. "Evolutionary Reinforcement Learning: A Systematic Review and Future Directions," Mathematics, MDPI, vol. 13(5), pages 1-33, March.
  • Handle: RePEc:gam:jmathe:v:13:y:2025:i:5:p:833-:d:1603779
    as

    Download full text from publisher

    File URL: https://www.mdpi.com/2227-7390/13/5/833/pdf
    Download Restriction: no

    File URL: https://www.mdpi.com/2227-7390/13/5/833/
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Zhang, Huizhen & An, Tianbo & Yan, Pingping & Hu, Kaipeng & An, Jinjin & Shi, Lijuan & Zhao, Jian & Wang, Jingrui, 2024. "Exploring cooperative evolution with tunable payoff’s loners using reinforcement learning," Chaos, Solitons & Fractals, Elsevier, vol. 178(C).
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Yongxiao Xie & Shian Song, 2025. "Deep Reinforcement Learning Method for Wireless Video Transmission Based on Large Deviations," Mathematics, MDPI, vol. 13(15), pages 1-18, July.
    2. Kyung-Soo Kim, 2025. "Utilization of Upper Confidence Bound Algorithms for Effective Subproblem Selection in Cooperative Coevolution Frameworks," Mathematics, MDPI, vol. 13(18), pages 1-36, September.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Wang, Jiaoyuan & Yang, Yanlong, 2026. "Research on the resilience of reputation mechanism to the cooperative environment in the face of external shocks," Applied Mathematics and Computation, Elsevier, vol. 510(C).
    2. Huang, Yijie, 2025. "The evolution of cooperation in multi-games with reinforcement learning," Chaos, Solitons & Fractals, Elsevier, vol. 201(P2).
    3. Zhang, Wei & Zhao, Dongkai & Jin, Xing & Zhang, Huizhen & An, Tianbo & Cui, Guanghai & Wang, Zhen, 2025. "Q-learning facilitates norm emergence in metanorm game model with topological structures," Chaos, Solitons & Fractals, Elsevier, vol. 195(C).
    4. An, Tianbo & Zhang, Huizhen & Zhang, Zhanshuo & Liu, Guanghui & Li, Jiayu & Chen, Liangyu & Wang, Zhen, 2025. "Cooperation dynamics driven by reinforcement learning with interactive diversity in structured populations," Chaos, Solitons & Fractals, Elsevier, vol. 201(P2).
    5. Wu, Binjie & Shen, Shaofei & Wang, Jiafeng & Wan, Haibin, 2025. "Q-learning promotes the evolution of fairness and generosity in the ultimatum game," Chaos, Solitons & Fractals, Elsevier, vol. 200(P2).
    6. Wang, Weining & Shang, Lihui & Wu, Yipeng & Hu, Mingjian & Wang, Weiyu, 2025. "Bio-inspired mechanism promotes cooperation in spatial public goods games," Chaos, Solitons & Fractals, Elsevier, vol. 200(P3).
    7. Yang, Yujin & Zhao, Dawei & Wang, Juan, 2025. "Evolution of cooperation in spatial public goods games driven by reinforcement learning and environmental feedback," Chaos, Solitons & Fractals, Elsevier, vol. 199(P1).
    8. Li, Yipeng & Hu, Xiangyue & Jin, Xing & Zhang, Huizhen & Yang, Jiajia & Wang, Zhen, 2025. "Environmental information perception enhances cooperation in stochastic public goods games via Q-learning," Applied Mathematics and Computation, Elsevier, vol. 504(C).
    9. Xie, Kai & Szolnoki, Attila, 2026. "Reinforcement learning in evolutionary game theory: A brief review of recent developments," Applied Mathematics and Computation, Elsevier, vol. 510(C).
    10. Qian, Yinuo & Zhao, Dawei & Xia, Chengyi, 2025. "Disease transmission in dynamic social networks constructed by reinforcement learning-driven preventive game," Chaos, Solitons & Fractals, Elsevier, vol. 199(P1).
    11. Shen, Shaofei & Zhang, Xuejun & Xu, Aobo & Duan, Taisen, 2024. "An adaptive exploration mechanism for Q-learning in spatial public goods games," Chaos, Solitons & Fractals, Elsevier, vol. 189(P1).
    12. Yan, Zeyuan & Zhao, Hui & Li, Li, 2025. "Dynamic role-switching in hypergraphs: Enhancing cooperation via adaptive punishment and reinforcement learning," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 677(C).
    13. Mangold, Gustavo C. & Vainstein, Mendeli H. & Fernandes, Heitor C.M., 2025. "Dilution, diffusion and symbiosis in spatial prisoner’s dilemma with reinforcement learning," Chaos, Solitons & Fractals, Elsevier, vol. 201(P3).
    14. Xie, Kai & Szolnoki, Attila, 2025. "Reputation in public goods cooperation under double Q-learning protocol," Chaos, Solitons & Fractals, Elsevier, vol. 196(C).
    15. Zheng, Guozhong & Zhang, Jiqiang & Deng, Shengfeng & Cai, Weiran & Chen, Li, 2024. "Evolution of cooperation in the public goods game with Q-learning," Chaos, Solitons & Fractals, Elsevier, vol. 188(C).
    16. Zhang, Yongqiang & Zheng, Zehao & Zhang, Xiaoming & Ma, Jinlong, 2025. "Dynamic punishment-reputation synergy drives cooperation in spatial public goods game," Applied Mathematics and Computation, Elsevier, vol. 506(C).
    17. Zhao, Jian & Zhang, Ran & An, Tianbo & Zhang, Huizhen & Tong, Daqun & Luo, Xun & An, Jinjin & Wang, Jingrui, 2025. "The effect of heterogeneous perceptions with environmental feedback on spatial social dilemmas," Chaos, Solitons & Fractals, Elsevier, vol. 192(C).

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jmathe:v:13:y:2025:i:5:p:833-:d:1603779. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address (email available below). General contact details of provider: https://www.mdpi.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.