Operational intelligence for SREs: treating incidents as reasoning problems
Alerts tell you something is wrong. Agents that recall past incidents, and a deterministic policy that decides how far they may act, close the gap to knowing what to do.
TL;DR
- The expensive part of an incident sits between the alert and the decision. More alerting does not shorten it.
- Give agents incident memory: retrieve similar past incidents and what fixed them, then reason over that record.
- Let memory earn autonomy, and let a deterministic policy outside the LLM decide how far an agent may act.
- Observe the reasoning, not only the outcome.
The costly minutes sit between the alert and the decision
Most reliability tooling is tuned for detection: more alerts, denser dashboards, faster paging. Little of it shortens the stretch that hurts, where you know something is wrong and nobody yet knows what to do. The README of the Agentic Reliability Framework (ARF) names it plainly: the real business loss happens between "Something is wrong" and "We know what to do."
That stretch is a reasoning problem. The on-call engineer is doing three jobs at once: retrieval (have we seen this before?), diagnosis (what is different this time?) and a risk call (is it safe to act?). An alert rule does none of them. By operational intelligence I mean agents that do the first two against your incident history and hand the third to a policy you wrote before the pager went off.
Treating an incident as a reasoning problem with memory
ARF splits the work across three agents. A Detection Agent watches telemetry and forecasts anomalies. A Recall Agent retrieves similar incidents, actions and outcomes from a RAG graph backed by FAISS. A Decision Agent applies deterministic policies and reasons over those historical outcomes.
Recall is the part that changes the job. Research on LLM multi-agent systems makes the general case: agents that reference past interactions with contextual similarity to the current query give more relevant and accurate responses. For an SRE, past interactions means postmortems, the actions taken and whether they worked.
An illustration, not a measured case: a latency alert fires on checkout. Recall returns two prior incidents with the same signature. One was fixed by rolling back a circuit-breaker config; in the other, the rollback did nothing because a slow dependency was the cause. The Decision Agent now has what a fresh prompt never has: a precedent and a counter-example to rule out before recommending the rollback.
Memory surfaces precedent; it does not prove cause. When the telemetry fits no precedent, causal graphs help, because they allow counterfactual reasoning about remediation options and their likely impact before anyone acts. That is a component you would add yourself; the ARF README does not describe one.
Let memory decide how much the agent may do
Recall also answers the question SREs worry about most: when may the agent act alone? Neel Shah's three tiers on DevOps.com give a workable answer. Tier 1 automates well-understood, reversible incidents. Tier 2 lets the agent investigate and recommend while a human approves. Tier 3 keeps people in command for novel, complex or high-impact incidents. The tier follows from familiarity, blast radius, reversibility and diagnostic confidence.
Two of those four are memory. Familiarity asks whether the pattern recurs; confidence rests in large part on whether the same action worked before. An agent with no incident history cannot earn Tier 1, which is the right outcome.
The article's firmest rule is that the final autonomy decision lives outside the LLM: the model can recommend a tier, and a deterministic policy enforces it. A starting policy fits in one function. The thresholds are mine; tune them against your own history.
from dataclasses import dataclass
@dataclass
class Proposal:
service: str
action: str
services_touched: int
reversible: bool
destructive: bool
past_matches: int # similar incidents the Recall Agent found
past_successes: int # of those, how many this action resolved
ALLOWLIST = {"orders-api", "search"}
def tier(p: Proposal) -> int:
if (p.service not in ALLOWLIST or p.destructive or not p.reversible
or p.services_touched > 1):
return 3
record = p.past_successes / p.past_matches if p.past_matches else 0.0
if p.past_matches >= 5 and record >= 0.9:
return 1
return 2
print(tier(Proposal("orders-api", "rollback", 1, True, False, 12, 12))) # 1
print(tier(Proposal("orders-api", "rollback", 1, True, False, 2, 2))) # 2
print(tier(Proposal("payments", "rollback", 1, True, False, 20, 19))) # 3The same rollback lands in three tiers depending on where it runs and how much history backs it. ARF draws a similar line in its own design: the OSS edition never executes changes, and the framework's safety layer caps blast radius at three services. That makes it a safe place to try recall with no write access anywhere.
Watch the reasoning, not only the outcome
A reasoning agent can fail quietly: right answer by luck, or a sound chain that ran out of room. Two signals catch that.
Score the trajectory. Datadog, building an evaluation platform for its SRE agent, chose to evaluate the agent's trajectory instead of scoring only the final conclusion. Replay past incidents and check which evidence the agent pulled, not just what it concluded.
Watch context pressure. A high rate of max_tokens finish reasons is a direct signal of a context-management problem, and for a recall agent my first suspect is that it retrieved too much. If you export finish reasons as a counter (the name agent_finish_reason_total is yours to choose), the rule is short. The 10% threshold is a starting guess.
groups:
- name: agent-reasoning
rules:
- alert: AgentMaxTokensRate
expr: |
sum by (agent) (rate(agent_finish_reason_total{reason="max_tokens"}[5m]))
/ sum by (agent) (rate(agent_finish_reason_total[5m])) > 0.1
for: 10mA rollout path that starts read-only
- Export your incident history and classify it by frequency, root cause, remediation, blast radius, reversibility and time to recovery. This is the memory; its quality caps everything downstream.
- Install ARF OSS in a sandbox with
pip install agentic-reliability-framework==3.3.9and replay old incidents through it. It runs entirely in memory with bounded retention, so treat it as an evaluation tool, not a system of record. - Pilot in shadow mode: the agent investigates real incidents without execution rights, and you compare its diagnosis and proposed action with what the human did.
- Put the tier policy in version control next to the runbooks. Promote an incident type to Tier 1 only when its record backs it.
- Re-score trajectories on replayed incidents with every prompt or model change, and alert on max_tokens rates the way you alert on latency.
What does not work: thin, inconsistent postmortems. Recall over noise returns noise, and no model fixes that. If your incident write-ups are three lines long, start there.
Sources
- GitHub - petterjuan/agentic-reliability-framework: ARF is an agentic reliability intelligence platform that separates decision intelligence (OSS) from governed execution (Enterprise), enabling autonomous operations with deterministic safety guarantees., GitHub
- LLM Multi-Agent Systems: Challenges and Open Problems, arxiv.org
- How Causal Reasoning Addresses the Limitations of LLMs in Observability, InfoQ
- How we built a real-world evaluation platform for autonomous SRE agents at scale, Datadog
- Site Reliability Engineering for AI Agent Systems: Observability, Incident Response, and Operational Patterns | Zylos Research, Zylos
sreaiincident-responseobservabilityagentic