ARIPAN

Technology & Consulting

Agentic Operations Platform for Autonomous Incident Response

A full platform architecture and a runnable reference implementation that proves the alert-to-root-cause agent loop end to end, with governance built in by construction.

01

The challenge

A global technology organisation needed an operating model for AI-driven operations across an estate of roughly 2,000 microservices spanning AWS, Azure, and GCP on multi-cluster Kubernetes. Its 24×7 NOC handled 300+ incidents a month against a four-hour MTTR, with the same root causes recurring. Leadership wanted autonomous triage, AI root-cause analysis, and auto-remediation — without handing production to an unsupervised model.

02

What Aripan built

  • An agent taxonomy with defined boundaries — triage, RCA, remediation, cost, and change-risk agents — with an explicit position on where the LLM reasons and where deterministic automation decides.
  • A working demo stack: a buggy target service, a Prometheus-driven alerter, a triage agent that converts alerts into structured incidents, and an RCA agent that puts every hypothesis through five causation gates before scoring confidence.
  • Grounding infrastructure: a Neo4j knowledge graph of services and deployments, a vector store for semantic recall of past incidents, PostgreSQL as the incident store, and an LLM gateway handling routing, JSON-schema validation, and audit logging.
  • Governance by construction: remediation is proposed against a bounded actuator catalogue and never auto-executed; every agent decision is written to an append-only audit trail.
03

Outcome

The client moved from an aspiration to an architecture with a demonstrable agent loop: alert to confidence-scored root cause in roughly 60–90 seconds in the reference environment, with full traceability. A separate proof of concept evaluated LLM incident classification against 2,600 log entries from six sources to establish accuracy, latency, and cost baselines before any platform investment.

Human-in-the-loop by design, with a full append-only audit trail


Client names are withheld by agreement. Details are kept generic to preserve anonymity.

← All case studies