Author
Abstract
Hyperscale data centers use switching fabrics to keep large numbers of computers continuously connected. However, with the growing demands of distributed AI workloads and edge deployments, relying only on passive fault tolerance is no longer enough to ensure network resilience at scale. Traditional redundancy designs handle isolated failures well, but they break down when gradual degradations accumulate silently across thousands of interconnected paths before any single alarm is triggered. This gap in visibility and response is a key challenge addressed by this paper. In contrast, a self-healing Network fabric continuously performs anomaly detection, correlates telemetry signals across multiple layers into clear diagnoses, and autonomously executes pre-validated mitigations. A controlled evaluation using an 8-leaf, 4-spine Clos simulation environment characterized detection latency, mitigation time, and convergence behavior across ten distinct failure classes and compared outcomes against a human-operator baseline; detection latency averaged 847 milliseconds, while automated mitigation time of 2.3 seconds represented a 19.6× mitigation time improvement over the human baseline with zero packet loss across 42 automated remediations during a 7-day simulation period. A four-layer governance model integrates blast-radius containment directly into the remediation pipeline, distinguishing the framework from intent-based and AIOps-driven approaches that treat governance separately. Incident knowledge repositories enable pattern recognition and policy retrieval at machine speed, while rollback mechanisms and full audit provenance ensure governance obligations are satisfied on every automated action. These results demonstrate that autonomous remediation can meet the operational demands of AI-scale infrastructure without compromising safety governance.
Suggested Citation
Vijaya Bhaskar Methuku, 2026.
"Self-Healing Datacenter Network Fabrics: Autonomous Remediation in Hyperscale and Distributed Cloud Infrastructure,"
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, International Journal of Scientific Research in Computer Science, Engineering and Information Technology, vol. 12(2), pages 616-632, April.
Handle:
RePEc:jbh:ijsrcs:v12:y2026:i2:id:1964
DOI: 10.32628/CSEIT26121392
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT26121392
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:jbh:ijsrcs:v12:y2026:i2:id:1964. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (USA) (email available below). General contact details of provider: https://ijsrcseit.com/home .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.