Author
Listed:
- Li, Jiaqi
- Wu, Jingda
- He, Hongwen
Abstract
Eco-driving at signalized intersections must balance energy efficiency, traffic efficiency, and safety under dynamic signal and traffic constraints. Although deep reinforcement learning (DRL) has shown promise for this task, policies trained in nominal environments may lose reliability when unseen signal timings or abrupt preceding-vehicle maneuvers expose experience-sparse boundary states. Under the conventional DRL train-and-test paradigm, such evaluation failures are usually recorded as terminal outcomes rather than reused for policy improvement. To address this missing data loop, this paper proposes a closed-loop offline policy improvement framework that recycles evaluation trajectories for successor-policy refinement. Specifically, safe and failure trajectories from a predecessor policy are collected into a mixed-quality offline dataset and reused through offline reinforcement learning, enabling value-guided improvement while regularizing policy updates toward data-supported actions. Results are evaluated under randomized signal-timing scenarios and an emergency-braking case. The proposed framework improves safety robustness over conventional DRL baselines and matches or exceeds the safety performance of a safety-constrained DRL baseline, with less conservative and less reactive behavior. The refined policy attains 6.65 ± 0.71 kWh/100 km, the lowest among all deployable methods, corresponding to 92.7% of the Dynamic Programming based theoretical reference computed under ideal information conditions. Vehicle-in-the-loop tests further indicate that smoother longitudinal behavior remains observable after real execution dynamics are introduced.
Suggested Citation
Li, Jiaqi & Wu, Jingda & He, Hongwen, 2026.
"Closed-loop offline policy improvement for eco-driving at signalized intersections via evaluation-trajectory recycling,"
Energy, Elsevier, vol. 360(C).
Handle:
RePEc:eee:energy:v:360:y:2026:i:c:s0360544226019444
DOI: 10.1016/j.energy.2026.141837
Download full text from publisher
As the access to this document is restricted, you may want to
for a different version of it.
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:energy:v:360:y:2026:i:c:s0360544226019444. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.journals.elsevier.com/energy .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.