Author
Listed:
- Rachman, Rifny
- Tingey, Josh
- Allmendinger, Richard
- Shukla, Pradyumn
- Pan, Wei
Abstract
This study develops a generalised multi-objective, multi-echelon supply chain optimisation model with non-stationary markets based on a Markov decision process, incorporating economic, environmental, and social considerations. The model is evaluated using a multi-objective reinforcement learning (RL) method, benchmarked against an originally single-objective RL algorithm modified with weighted sum using predefined weights, and a multi-objective evolutionary algorithm (MOEA)-based approach. We conduct experiments on varying network complexities, mimicking typical real-world challenges using a customisable simulator. The model determines production and delivery quantities across supply chain routes to achieve near-optimal trade-offs between competing objectives, approximating Pareto front sets. The results demonstrate that the primary approach provides the most balanced trade-off between optimality, diversity, and density, further enhanced with a shared experience buffer that allows knowledge transfer among policies. In complex settings, it achieves approximately ten times higher hypervolume than the MOEA-based method and generates solutions that are substantially denser than those produced by the modified single-objective RL method. Moreover, it ensures stable production and inventory levels while minimising demand loss.
Suggested Citation
Rachman, Rifny & Tingey, Josh & Allmendinger, Richard & Shukla, Pradyumn & Pan, Wei, 2026.
"Reinforcement learning for multi-objective multi-echelon supply chain optimisation,"
European Journal of Operational Research, Elsevier, vol. 334(3), pages 942-962.
Handle:
RePEc:eee:ejores:v:334:y:2026:i:3:p:942-962
DOI: 10.1016/j.ejor.2026.02.002
Download full text from publisher
As the access to this document is restricted, you may want to
for a different version of it.
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:ejores:v:334:y:2026:i:3:p:942-962. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/eor .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.