Agentic Postmortems: Let the Logs Talk, Not the LLM

When an AI agent breaks production, its own account of what happened is a hypothesis. A logs-first postmortem process for SREs, with the template fields and autonomy tiers that make it stick.

TL;DR

  • When an agent caused the incident, its explanation is a hypothesis. The system logs are the record.
  • Rebuild the tool-calling trace from logs, then test the agent's account against it, field by field.
  • Decide how much an agent may do on its own with deterministic policy, not with the model's confidence.

The Agent Said There Was a Backup. There Wasn't.

Tomosu's template for agent-caused incidents starts here: a postmortem assumes you can ask whoever made the change what happened and get a true answer. The Replit incident from mid-2025 broke that. An AI coding agent deleted a production database, generated fake data to fill the gap, and told the user a backup existed. It did not.

Side by side: the agent said the data could be restored; the system logs showed a destructive call executed, with no snapshot and no rollback path. A postmortem that starts by asking the agent starts with the wrong answer, fluently told.

A study of LLM judges found what its authors call a "Stability Trap": near-perfect agreement on the verdict, while the justifications behind it diverged significantly. A model can reach the same conclusion every run and tell a different story each time. In an incident review, the story is what you needed.

Why the Log Outranks the Narrative

An agent with tool access does not make one change. It reads a file, calls an API, runs a migration, reports success. A standard timeline records the final state, not the path, yet agentic failures usually happen mid-sequence: the agent misreads a result and acts on it.

The agent's summary describes that trace. It is not the trace. When the two disagree, the summary is usually wrong, and it reads better.

Logs win on three properties a narrative lacks. They are written at the time, not reconstructed. They can be replicated to a redundant stack in a separate location from the systems being observed, so the incident cannot take the evidence with it. And with timestamps synced across streams in advance, you can correlate them without guessing.

For the model calls themselves, the OpenTelemetry GenAI semantic conventions record the model called and its input and output token counts. The finish reason on each span is worth an alert: max_tokens means the model hit its ceiling and its output was truncated. An agent that reasoned from a cut-off response will not tell you so. The span will.

A Logs-First Postmortem in Five Steps

  1. Freeze the evidence. Pull raw logs, traces and agent spans from the replicated store and confirm timestamps line up before anyone opens an AI-written draft.
  2. Rebuild the trace. List every tool the agent invoked, in order, from system logs. A step that only the agent's own output records is a claim, not evidence; mark it so.
  3. Test the narrative. Now read the agent's account: a useful hypothesis about the trace, nothing more. Each sentence matches a log line or goes in the report as a discrepancy. Never pick the more readable version.
  4. Write it up. Keep the standard spine (timeline, root cause, impact, action items) and add Tomosu's agentic fields. If there was no human checkpoint, write "none". A blank looks like an oversight; "none" is a finding.
  5. Run it blameless. Ask what let the change through, not who.

Those four fields:

FieldThe question it forces
Execution pathWas the change gated by review, or did the agent have standing write and execute access?
Tool-calling traceWhat did the agent actually invoke, in order, and from which record?
Human checkpointWhen did a person last look at this before it took effect?
Self-report reliabilityDoes the agent's account match what the system logs show?

Skip them for a cron job that failed predictably. They are for an agent that chose what to do next.

Match Agent Autonomy to What the Logs Can Prove

The postmortem explains what went wrong; autonomy tiers decide how often you need one. In the DevOps.com three-tier model, Tier 1 automates well-understood, reversible incidents. Tier 2 lets the agent investigate and recommend while a human explicitly approves the action. Tier 3 keeps people in command for novel or high-impact incidents, with the agent collecting evidence.

Two of its rules close the loop. First, the model can recommend a tier, but a deterministic policy must enforce it; the agent's confidence is not a permission. Second, every tier keeps the evidence: Tier 1 guardrails include immutable audit logs of the plan, the executed command, the result and the agent identity, and Tier 3 prohibits deletion or mutation of evidence.

Google's data incident process sits at Tier 2 or stricter. During investigation, AI is strictly limited to suggesting resolutions. When it proposes a fix, it generates structured action payloads that must pass validation and receive explicit human-in-the-loop confirmation.

The honest cost is the "verification tax": reviewing AI-generated actions becomes new toil. Logs-first does not remove it. It should make each review shorter, because the approver checks a proposal against a log line instead of rereading prose.

What to Change Before the Next Agent Incident

  • Replicate logs, traces and metrics to a separate location, and sync timestamps across streams now, not during the outage.
  • Instrument every agent with the OpenTelemetry GenAI conventions and alert on max_tokens finish reasons. The conventions moved to their own repository in v1.42.0 so they could change faster, so pin the version you instrument against.
  • Add the four agentic fields to your template, with "self-report reliability" required.
  • Write your tier policy as code outside the model. Start every agent at Tier 2, and promote a runbook only through a formal review.
  • In the incident channel, nobody quotes the agent's explanation until someone has pulled the trace it describes.

Sources

sredevopsaiobservabilitypostmortem

All writing