Loading patterns…
Agentic SRE (Self-Healing Operations)(ASRE)
A closed-loop operations architecture in which agents keep an external system healthy: detect an anomaly, diagnose the probable root cause, execute a policy-bounded remediation, then verify recovery against reliability objectives before closing the loop. It is typically built as a role-specialized team (detector, diagnoser, remediator, verifier) running under human-on-the-loop governance, where engineers define policy, guardrails, and the set of allowed actions while the agents execute within those bounds and report. Least-privilege permissions and policy-as-code keep remediations inside a safe envelope, with higher-risk actions gated on human approval. Distinct from `predictive-agent-fault-tolerance`: that keeps the agent system itself alive, whereas this is agents performing reliability work on the external systems they operate.
In 30 seconds
- What
- Agents detect anomalies, diagnose root causes, execute policy-bounded fixes, then verify recovery against SLOs in a closed loop under human oversight.
- When to use
- External systems need autonomous recovery from known failure modes while keeping humans in control of what actions agents can take.
- Watch out
- Overly permissive policies or weak verification can cause agents to mask symptoms instead of fixing root causes or cascade failures.
Ask the AI expert about this pattern
Opens the assistant with your question prefilled. You review it before sending.
Agentic SRE (Self-Healing Operations): Overview
A closed-loop operations architecture in which agents keep an external system healthy: detect an anomaly, diagnose the probable root cause, execute a policy-bounded remediation, then verify recovery against reliability objectives before closing the loop. It is typically built as a role-specialized team (detector, diagnoser, remediator, verifier) running under human-on-the-loop governance, where engineers define policy, guardrails, and the set of allowed actions while the agents execute within those bounds and report. Least-privilege permissions and policy-as-code keep remediations inside a safe envelope, with higher-risk actions gated on human approval. Distinct from `predictive-agent-fault-tolerance`: that keeps the agent system itself alive, whereas this is agents performing reliability work on the external systems they operate.
- Closed loop: detect, diagnose, remediate, verify
- Role-specialized team: detector, diagnoser, remediator, verifier
- Policy-bounded autonomy with a least-privilege set of allowed actions
- Human-on-the-loop: engineers set policy, agents execute and report
- Recovery verified against SLOs before an incident is closed
- Automatic rollback and generated incident reports for audit
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.
References
The papers, specifications, and repositories this pattern is based on.
From the engineer behind this catalog
Get your agent architecture reviewed
This page documents one pattern. Your system runs dozens, and most failures live in how they fit together. Have the whole design reviewed against the 288 patterns in this catalog: architecture, reliability, evaluation and cost, every finding mapped to the pattern that fixes it.
€750 instead of €1,500, one week, written report and walkthrough call, until 30 September