Article Information

Authors Nagaraja. K.V.
Article Type Research Article
Language English
Journal North Asian International Research Journal of Sciences, Engineering & I.T.
ISSN 2454-7514
Volume 11
Issue 5
Pages 1-9
Publication Year 2025
Publication Date November 01, 2025
DOI URL https://doiglobal.org/10.2025/NAIRJCSEIT.005

Abstract

As enterprise applications transition toward dynamic microservice meshes, classical observability pipelines remain heavily reliant on human-in-the-loop intervention for post-anomaly remediation. While contemporary predictive models successfully detect runtime degraded states, automating recovery actions—such as (dynamic traffic rerouting, selective pod isolation, and automated circuit breaking—without triggering unstable control loops remains a critical challenge. This paper introduces ResilNet, an autonomous, zero-touch remediation framework that combines dynamic Graph Attention Networks (GAT) with Deep Q-Networks (DQN). ResilNet continuously ingests real-time call topologies and telemetry streams, constructing state-space representations that model multi-hop cascading dependencies. By mapping continuous infrastructure states into a constrained, safety-bounded action space, the framework learns optimal mitigation policies that minimize Mean Time to Recovery MTTR) while preventing secondary operational degradation. Evaluated across a 1,000-node Kubernetes cluster under complex chaotic fault injections, ResilNet achieved an average MTTR reduction of 68.4% compared to automated rule-based operators and reduced system-wide service disruption during transient failures to under 1.2%.

Keywords

Microservices Remediation Deep Reinforcement Learning Graph Attention Networks

60

Total Views

0

Total Downloads

2025

Publication Year

Active

DOI Status

DOI Active Record

This scholarly article has been successfully registered and assigned a persistent DOI through DOIGLOBAL.

DOI: 10.2025/NAIRJCSEIT.005

References

1. Brockman, G., et al. (2016). OpenAI Gym. arXiv preprint arXiv:1606.01540.
2. Mnih, V., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529-533.
3. Veličković, P., Cucurull, G., Casanova, A., Romero, A., Liò, P., & Bengio, Y. (2018). Graph Attention Networks. International Conference on Learning Representations (ICLR).
4. Zhou, X., et al. (2021). Fault Analysis and Automated Recovery in Cloud-Native Architectures. IEEE Transactions on Software Engineering,

Research Tools

Share This Research

Recommended Citation

Nagaraja. K.V. (2025). Zero-Touch Resilience: Autonomous Fault Remediation in Distributed Microservices via Graph-Attentive Deep Q-Learning. North Asian International Research Journal of Sciences, Engineering & I.T.. DOI: https://doiglobal.org/10.2025/NAIRJCSEIT.005

Publisher Information

Journal:
North Asian International Research Journal of Sciences, Engineering & I.T.

DOI Provider:
DOIGLOBAL

This scholarly record is permanently registered through DOIGLOBAL DOI Infrastructure and remains accessible using its persistent DOI identifier.

💬
DOI Global Assistant
👋 Welcome to DOI Global Assistant

I can help you with:

🔍 DOI Search
📚 DOI Registration
👥 Membership Information
🔐 Publisher Login
🛠 Publisher Services
📄 DOI Certificates
📞 Contact Support

Popular Questions:

• What is a DOI?
• What does DOI stand for?
• How do I register a DOI?
• Membership benefits
• Publisher login

Example DOI:

10.2026/JOURNAL.006

💡 You can search by:
DOI Number, Article Title, Author Name, or Journal Name.