IDEAS home Printed from https://ideas.repec.org/a/gam/jftint/v18y2026i7p337-d1975492.html

Regional Strategy Composition: A Hierarchical-Action Reinforcement Learning Framework for Dynamic Smart-Meter Association over 5G NR mMTC Networks

Author

Listed:
  • Muhammed Al-Ali

    (Department of Computer Science and Engineering, Qatar University, Doha P.O. Box 2713, Qatar)

  • Esteban Inga

    (Smart Grids Research Group (GIREI), Universidad Politécnica Salesiana, Cuenca 010102, Ecuador)

  • Juan Inga

    (Telecommunications and Telematic Research Group (GITEL), Universidad Politécnica Salesiana, Cuenca 010102, Ecuador)

  • Elias Yaacoub

    (Department of Computer Science and Engineering, Qatar University, Doha P.O. Box 2713, Qatar)

Abstract

Advanced Metering Infrastructure (AMI) over 5G New Radio (NR) massive machine-type communication (mMTC) networks require efficient and adaptive communication mechanisms to support reliable data delivery for large numbers of smart meters under dynamic traffic and channel conditions. In this work, we propose a framework in which each smart meter chooses, at runtime, whether to transmit directly to the base station (BS) or via a nearby Data Aggregation Point (DAP). The optimal choice is dynamic and depends on DAP buffer occupancy, periodic congestion, channel quality, and packet deadline pressure. Formulating this as a per-meter binary decision yields an action space of size 2 N for N meters, which is intractable for reinforcement learning (RL). We reformulate the problem as regional strategy composition: the RL agent selects one parameterized association strategy for each DAP region from a small library of interpretable rules, and a deterministic mapping expands the regional choice into per-meter modes. It reduces the policy action space from 2 N to K D , where D is the number of DAPs and K the number of strategies, while preserving meter-level control granularity. We evaluate Proximal Policy Optimization (PPO) and Deep Q-Network (DQN) controllers against eight meter-level baselines on a 5G NR-calibrated simulator with 1500 m, six DAPs, deadline-bounded delivery, stale channel-state information, and phase-offset congestion cycles. Across three traffic regimes and five random seeds, PPO improves packet delivery ratio (PDR) over the strongest heuristic by +0.63, +2.41, and +2.66 percentage points under baseline, high-load, and bursty-cycle conditions, respectively; all gains are statistically significant (paired t -test, p < 0.001 ; Cohen’s d up to 5.12), and the advantage grows with traffic stress. The results show that learned regional composition of classical heuristics outperforms any single fixed heuristic precisely when no individual rule is globally optimal.

Suggested Citation

  • Muhammed Al-Ali & Esteban Inga & Juan Inga & Elias Yaacoub, 2026. "Regional Strategy Composition: A Hierarchical-Action Reinforcement Learning Framework for Dynamic Smart-Meter Association over 5G NR mMTC Networks," Future Internet, MDPI, vol. 18(7), pages 1-24, June.
  • Handle: RePEc:gam:jftint:v:18:y:2026:i:7:p:337-:d:1975492
    as

    Download full text from publisher

    File URL: https://www.mdpi.com/1999-5903/18/7/337/pdf
    Download Restriction: no

    File URL: https://www.mdpi.com/1999-5903/18/7/337/
    Download Restriction: no
    ---><---

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jftint:v:18:y:2026:i:7:p:337-:d:1975492. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address (email available below). General contact details of provider: https://www.mdpi.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.