Author
Listed:
- Prudvi Saisaran Ponduru
- Pavani Priya Vyshnavi Nandanavanam
- Sai Kesav Kumar Ponduru
Abstract
Cloud infrastructure failures are increasingly difficult to detect, diagnose, and remediate because production environments combine microservices, Kubernetes control loops, service meshes, serverless workloads, infrastructure-as-code, continuous delivery, and heterogeneous telemetry. Reactive monitoring and manual incident response remain necessary, but they do not scale to the volume, velocity, and causal complexity of modern cloud operations. This paper provides a structured synthesis of scholarly, industry, and standards-based work on AI-driven self-healing for cloud infrastructure, with emphasis on AIOps, AgentOps, LLM-based cloud operations, anomaly detection, causal root cause analysis, graph learning, reinforcement learning, automated remediation, Kubernetes self-healing, observability, chaos engineering, and self-healing infrastructure-as-code. We propose SAFE-HealCloud, a safety-aware, agentic, feedback-driven framework that integrates telemetry ingestion, multimodal observability, anomaly detection, failure prediction, causal RCA, retrieval-augmented LLM reasoning, policy guardrails, risk-scored remediation planning, controlled execution, verification, rollback, human approval, and continuous learning. A formal model defines cloud state, observability vectors, failure states, action spaces, remediation policies, rewards, constraints, and reliability objectives. We also specify reproducible Kubernetes-based evaluation designs, metrics, algorithms, risk controls, and operational use cases. The central finding is that near-term practical value lies in graduated autonomy: low-risk reversible actions can be automated, medium-risk actions should be canaried and policy-gated, and high-risk changes should remain human-approved. Designed in this way, AI-driven self-healing can reduce detection and recovery times, preserve error budgets, improve operator productivity, and strengthen digital resilience without sacrificing safety, auditability, or governance.
Suggested Citation
Prudvi Saisaran Ponduru & Pavani Priya Vyshnavi Nandanavanam & Sai Kesav Kumar Ponduru, 2026.
"SAFE-HealCloud: Safety-Aware, Agentic Self-Healing for Cloud Infrastructure,"
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, International Journal of Scientific Research in Computer Science, Engineering and Information Technology, vol. 12(4), pages 163-188, July.
Handle:
RePEc:jbh:ijsrcs:v12:y2026:i4:id:2121
DOI: 10.32628/CSEIT26124220
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT26124220
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:jbh:ijsrcs:v12:y2026:i4:id:2121. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (USA) (email available below). General contact details of provider: https://ijsrcseit.com/home .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.