Author
Listed:
- Chinatu M. Anyanwu
(Faculty of Computing, Maduka University, Ekwegbe-Enugu State)
- Nkiru C. Ogbonna
(Department of ICT/Innovation Centre, University of Nigeria, Nsukka)
- Mary Ofuru Kam
(Department of Computer Science, Veritas University, Bwari, Abuja, Nigeria)
- Stephen Uche Udeh
(Department of Computer Science, University of Nigeria, Nsukka)
- Ogechi Gift Onyedi
(School of Health Science, Maduka University, Ekwegbe-Enugu State)
Abstract
Personalized insulin dosing for Type 1 diabetes mellitus (T1DM) remains challenging because of complex glucose–insulin dynamics and substantial patient variability. Reinforcement learning (RL) has emerged as a promising approach for adaptive insulin management, yet the reliability of learned policies depends heavily on reward design and evaluation strategy. This study compares three actor–critic RL algorithms: Soft Actor-Critic (SAC), Advantage Actor-Critic (A2C), and Proximal Policy Optimization (PPO) for personalized insulin dosing using real-world continuous glucose monitoring, insulin delivery, basal insulin, and meal intake data from the OhioT1DM dataset. A custom Gymnasium-based environment was developed, and all algorithms were trained under identical conditions for 100,000 timesteps. Performance was evaluated using cumulative reward together with clinically relevant measures, including Time in Range (TIR) and insulin dosing behaviour. Although A2C and PPO achieved higher cumulative rewards than SAC, both converged to near-zero insulin dosing policies that exploited the reward formulation rather than learning clinically meaningful glucose regulation. In contrast, SAC maintained adaptive dosing behaviour, achieving a TIR of 72.71% with an average insulin dose of 1.769 U/step. These findings show that higher cumulative reward does not necessarily correspond to better clinical decision-making in open-loop reinforcement learning environments. The study highlights the importance of behaviour-focused evaluation alongside conventional reward metrics and provides practical insights for developing safer and more reliable reinforcement learning systems for personalized diabetes management.
Suggested Citation
Chinatu M. Anyanwu & Nkiru C. Ogbonna & Mary Ofuru Kam & Stephen Uche Udeh & Ogechi Gift Onyedi, 2026.
"Reinforcement Learning for Personalized Insulin Dosing: A Comparative Study of A2C, SAC and PPO on Real-World Clinical Data,"
International Journal of Latest Technology in Engineering, Management & Applied Science, RSIS International, vol. 15(7), pages 1094-1104, August.
Handle:
RePEc:bjf:ijltem:v:15:y:2026:i:7:a:92
DOI: 10.51583/IJLTEMAS.2026.150700087
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bjf:ijltem:v:15:y:2026:i:7:a:92. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Dr. Pawan Verma (email available below). General contact details of provider: https://www.ijltemas.in/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.