Operational intelligence for SREs: treating incidents as reasoning problems

Alerts tell you something is wrong. Agents that recall past incidents, and a deterministic policy that decides how far they may act, close the gap to knowing what to do.

TL;DR

  • The expensive part of an incident sits between the alert and the decision. More alerting does not shorten it.
  • Give agents incident memory: retrieve similar past incidents and what fixed them, then reason over that record.
  • Let memory earn autonomy, and let a deterministic policy outside the LLM decide how far an agent may act.
  • Observe the reasoning, not only the outcome.

The costly minutes sit between the alert and the decision

Most reliability tooling is tuned for detection: more alerts, denser dashboards, faster paging. Little of it shortens the stretch that hurts, where you know something is wrong and nobody yet knows what to do. The README of the Agentic Reliability Framework (ARF) names it plainly: the real business loss happens between "Something is wrong" and "We know what to do."

That stretch is a reasoning problem. The on-call engineer is doing three jobs at once: retrieval (have we seen this before?), diagnosis (what is different this time?) and a risk call (is it safe to act?). An alert rule does none of them. By operational intelligence I mean agents that do the first two against your incident history and hand the third to a policy you wrote before the pager went off.

Treating an incident as a reasoning problem with memory

ARF splits the work across three agents. A Detection Agent watches telemetry and forecasts anomalies. A Recall Agent retrieves similar incidents, actions and outcomes from a RAG graph backed by FAISS. A Decision Agent applies deterministic policies and reasons over those historical outcomes.

Recall is the part that changes the job. Research on LLM multi-agent systems makes the general case: agents that reference past interactions with contextual similarity to the current query give more relevant and accurate responses. For an SRE, past interactions means postmortems, the actions taken and whether they worked.

An illustration, not a measured case: a latency alert fires on checkout. Recall returns two prior incidents with the same signature. One was fixed by rolling back a circuit-breaker config; in the other, the rollback did nothing because a slow dependency was the cause. The Decision Agent now has what a fresh prompt never has: a precedent and a counter-example to rule out before recommending the rollback.

Memory surfaces precedent; it does not prove cause. When the telemetry fits no precedent, causal graphs help, because they allow counterfactual reasoning about remediation options and their likely impact before anyone acts. That is a component you would add yourself; the ARF README does not describe one.

Let memory decide how much the agent may do

Recall also answers the question SREs worry about most: when may the agent act alone? Neel Shah's three tiers on DevOps.com give a workable answer. Tier 1 automates well-understood, reversible incidents. Tier 2 lets the agent investigate and recommend while a human approves. Tier 3 keeps people in command for novel, complex or high-impact incidents. The tier follows from familiarity, blast radius, reversibility and diagnostic confidence.

Two of those four are memory. Familiarity asks whether the pattern recurs; confidence rests in large part on whether the same action worked before. An agent with no incident history cannot earn Tier 1, which is the right outcome.

The article's firmest rule is that the final autonomy decision lives outside the LLM: the model can recommend a tier, and a deterministic policy enforces it. A starting policy fits in one function. The thresholds are mine; tune them against your own history.

from dataclasses import dataclass
 
@dataclass
class Proposal:
    service: str
    action: str
    services_touched: int
    reversible: bool
    destructive: bool
    past_matches: int    # similar incidents the Recall Agent found
    past_successes: int  # of those, how many this action resolved
 
ALLOWLIST = {"orders-api", "search"}
 
def tier(p: Proposal) -> int:
    if (p.service not in ALLOWLIST or p.destructive or not p.reversible
            or p.services_touched > 1):
        return 3
    record = p.past_successes / p.past_matches if p.past_matches else 0.0
    if p.past_matches >= 5 and record >= 0.9:
        return 1
    return 2
 
print(tier(Proposal("orders-api", "rollback", 1, True, False, 12, 12)))  # 1
print(tier(Proposal("orders-api", "rollback", 1, True, False, 2, 2)))    # 2
print(tier(Proposal("payments", "rollback", 1, True, False, 20, 19)))    # 3

The same rollback lands in three tiers depending on where it runs and how much history backs it. ARF draws a similar line in its own design: the OSS edition never executes changes, and the framework's safety layer caps blast radius at three services. That makes it a safe place to try recall with no write access anywhere.

Watch the reasoning, not only the outcome

A reasoning agent can fail quietly: right answer by luck, or a sound chain that ran out of room. Two signals catch that.

Score the trajectory. Datadog, building an evaluation platform for its SRE agent, chose to evaluate the agent's trajectory instead of scoring only the final conclusion. Replay past incidents and check which evidence the agent pulled, not just what it concluded.

Watch context pressure. A high rate of max_tokens finish reasons is a direct signal of a context-management problem, and for a recall agent my first suspect is that it retrieved too much. If you export finish reasons as a counter (the name agent_finish_reason_total is yours to choose), the rule is short. The 10% threshold is a starting guess.

groups:
  - name: agent-reasoning
    rules:
      - alert: AgentMaxTokensRate
        expr: |
          sum by (agent) (rate(agent_finish_reason_total{reason="max_tokens"}[5m]))
            / sum by (agent) (rate(agent_finish_reason_total[5m])) > 0.1
        for: 10m

A rollout path that starts read-only

  1. Export your incident history and classify it by frequency, root cause, remediation, blast radius, reversibility and time to recovery. This is the memory; its quality caps everything downstream.
  2. Install ARF OSS in a sandbox with pip install agentic-reliability-framework==3.3.9 and replay old incidents through it. It runs entirely in memory with bounded retention, so treat it as an evaluation tool, not a system of record.
  3. Pilot in shadow mode: the agent investigates real incidents without execution rights, and you compare its diagnosis and proposed action with what the human did.
  4. Put the tier policy in version control next to the runbooks. Promote an incident type to Tier 1 only when its record backs it.
  5. Re-score trajectories on replayed incidents with every prompt or model change, and alert on max_tokens rates the way you alert on latency.

What does not work: thin, inconsistent postmortems. Recall over noise returns noise, and no model fixes that. If your incident write-ups are three lines long, start there.

Sources

sreaiincident-responseobservabilityagentic

All writing