Agentic coding shifts the bottleneck to review, so redesign your CI/CD

AI-generated code has moved the slowest step in delivery from writing to reviewing. Here is how to rebuild the pipeline so its controls hold at agent speed, with a working circuit breaker and a 90-day rollout.

TL;DR

  • Review, not writing, is now the constraint. Move review policy into required pipeline checks.
  • Protect main, gate on agent review, and give agents tiered permissions and a circuit breaker.
  • Roll out over 90 days; code changes keep a human sign-off.

The bottleneck is no longer code creation

CloudBees surveyed 213 enterprise technology leaders for its 2026 State of Code Abundance Report. AI now writes or assists 61% of the code in the average enterprise. Asked where delivery is slowest, 57% of leaders point to reviewing, testing and deploying, and only 35% to writing.

Pipelines were built for writing to be the slow step, and the failures show it. 81% of leaders report more production issues they trace to AI-generated code, while 92% say they trust that code before it ships. Each change looks fine alone. Nobody can see how thousands of them behave together.

Teams have answered with more of the same. 70% say maintaining the test suite is now a bigger burden than writing code, and 97% changed their testing approach, most by running more tests.

The telling number: 93% have a formal process for reviewing and releasing AI-generated code, but only 56% say it is always enforced. A process a person has to remember fails once changes outnumber people. A control that is a required pipeline check is enforced on every change. That is the redesign.

What a governed agentic pipeline looks like

GitHub puts it well: the developer's job is shifting from writing code to owning the delivery system around it. Its enterprise roundup argues that the advantage goes to teams that run agents inside governed pipelines, combining automation with deterministic controls such as testing, security scans, branch protections and human review. Deterministic means the check gives the same answer every time, whoever or whatever wrote the change.

Four ideas make that concrete.

  • Maker-checker loop. One agent (the maker) proposes a change. A second agent (the checker) evaluates it against defined criteria and sends it back if it falls short. For example, a review agent checks a coding agent's PR against your security rules before a human sees it.
  • Tiered autonomy. Tier 1 is read-only observability and needs no approval. Tier 2 is low-risk remediation, such as retries and restarts, auto-approved and logged. Tier 3 is code changes and production deployments, and it always needs a human sign-off on a pull request or in Slack.
  • Policy layer. Rules for what the agent may read, suggest or execute, checked before it acts. In practice this is a policy engine such as Open Policy Agent sitting in front of Tier 2 and Tier 3 actions.
  • Model Context Protocol (MCP). MCP has become the main way coding agents connect to CI/CD systems, and agents invoke pipeline operations as MCP tool calls. The tools you expose over MCP are what bounds an agent, so expose few.

Concrete changes you can make to your CI/CD stack today

Do these in order; the first is the cheapest.

  1. Protect main. Branch protection rules on main stop an agent from merging its own PRs. It is a one-time setting that makes every other gate mandatory.
  2. Add an agent review gate. An AI reviewer flags security vulnerabilities, departures from your architectural conventions, and code that works but does not match the code around it. GitHub's Copilot code review is one such reviewer. Budget for the cost: an agent review step adds 45–120 seconds to every pipeline run.
  3. Run each agent in a clean environment. Every agent run should start from a fresh, reproducible environment with nothing shared between concurrent runs, so one bad run cannot contaminate the next.
  4. Select tests, but keep the full suite. Predictive test selection trains on which tests historically fail when given files change, and runs only the high-risk subset on each commit. Run the full suite nightly so the model's misses surface within a day.
  5. Add a circuit breaker. If an agent takes three or more actions within a rolling five-minute window without a successful outcome, pause it and page on-call.

The agent appends one JSON line per action; the breaker runs before every Tier 2 action.

# breaker.py: three or more actions in five minutes, none successful, trips the breaker.
import json
import sys
import time
 
WINDOW_SECONDS = 300
MIN_ACTIONS = 3
 
def should_pause(events, now):
    """events: [{"ts": epoch_seconds, "ok": bool},...] for one agent, oldest first."""
    recent = [e for e in events if now - e["ts"] <= WINDOW_SECONDS]
    return len(recent) >= MIN_ACTIONS and not any(e["ok"] for e in recent)
 
if __name__ == "__main__":
    path = sys.argv[1] if len(sys.argv) > 1 else "agent-actions.jsonl"
    with open(path) as f:
        events = [json.loads(line) for line in f if line.strip()]
    now = float(sys.argv[2]) if len(sys.argv) > 2 else time.time()
    if should_pause(events, now):
        print("circuit open: pausing agent, paging on-call")
        sys.exit(1)
    print("circuit closed")
$ printf '%s\n' '{"ts": 1000, "ok": false}' '{"ts": 1090, "ok": false}' '{"ts": 1200, "ok": false}' > agent-actions.jsonl
$ python3 breaker.py agent-actions.jsonl 1250; echo "exit $?"
circuit open: pausing agent, paging on-call
exit 1

GitHub Actions has no pause button, so the non-zero exit is the pause: it fails the job before the agent's next action runs. Wire that failure to your paging tool. Keep the action log outside the runner, in a small database table, because each job starts on a fresh machine.

Next steps for engineering leaders

A team of 10 to 50 engineers typically needs about 90 days to go from nothing to bounded autonomy.

  • Days 1–30, instrument. Add OpenTelemetry spans to every pipeline stage and build a history of runs: stage, duration, outcome, failure message.
  • Days 31–60, observe. Run the agent in observer mode. It posts a root-cause summary of every failure to Slack and changes nothing. Write down your Tier 1, 2 and 3 boundaries.
  • Days 61–90, bounded autonomy. Turn on Tier 1 and Tier 2 actions with full audit logging and the breaker in front of Tier 2, then predictive test selection. Tier 3 stays with humans.

Measure deployment frequency, MTTR, change failure rate and pipeline cost against the baseline from month one. Researchers also propose agent-specific KPIs: agent acceptance rate, supervision burden, escalation quality and prompt stability. They give no formulas, so here is how I would count the first two. Acceptance rate is agent PRs merged with no human commits on them, divided by agent PRs opened. Supervision burden is reviewer time per merged agent PR; the gap between review request and approval is a rough upper bound.

When not to do this: if your team cannot yet say who owns a failed deploy, fix that first. In the CloudBees survey, only 12% of organizations have a team responsible for AI-generated code. Agents in a pipeline nobody owns just produce failures faster.

Sources

ci/cdaidevopsautomationsoftware-engineering

All writing