Author
Listed:
- Zihan Xu
(School of Computer Science and Technology, Zhejiang Sci-Tech University, Hangzhou 310018, China)
- Chengqun Wang
(School of Computer Science and Technology, Zhejiang Sci-Tech University, Hangzhou 310018, China
Zhejiang Engineering Research Center of Industrial Internet Communication Technology, Zhejiang Sci-Tech University, Hangzhou 310018, China
Zhejiang Key Laboratory of Digital Fashion and Data Governance, Zhejiang Sci-Tech University, Hangzhou 310018, China)
Abstract
Future Internet applications such as intelligent transportation, immersive services, and edge-assisted artificial intelligence require latency-sensitive service provisioning at the network edge. In containerized mobile edge computing (MEC), service orchestration is not only a task-offloading problem, but also a task–container–image constrained decision problem: an offloaded task can be executed only when the required runtime container is active, and a newly activated container must be supported by a locally cached service image. This dependency couples task placement, runtime container caching, and persistent image caching under limited RAM and ROM resources. To address this challenge, this paper proposes HAM-MADDPG, a dependency-aware hierarchical action-masked multi-agent reinforcement learning algorithm for joint task offloading and image–container caching in containerized MEC networks. HAM-MADDPG decomposes the monolithic orchestration decision into three causally ordered policy layers: task offloading, runtime container caching, and persistent image caching. Each layer learns a structured subproblem conditioned on upstream realized decisions, while dynamic action masking and feasibility-aware action realization guide the learned policies toward executable decisions satisfying task–container and container–image constraints. Extensive simulations under dynamic service demands and heterogeneous edge resources show that HAM-MADDPG achieves more stable convergence than non-hierarchical reinforcement learning baselines and reduces long-term system latency by approximately 14–25% compared with representative heuristic and flat DRL baselines.
Suggested Citation
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jftint:v:18:y:2026:i:6:p:315-:d:1963810. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address
(email available below). General contact details of provider: https://www.mdpi.com .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.