Engineering Reliable Agentic Loops in Production
Iteration limits, idempotent tool calls, the cost-accuracy numbers behind four orchestration patterns, queue-depth autoscaling and pass^k, with runnable Python for each guardrail.
TL;DR
- Give every loop three hard stops: an iteration cap, a repetition check and a token budget.
- Key every side-effecting tool call so a retry replays the first result instead of acting twice.
- Default to one agent; reach for a hierarchical supervisor when you need several, and put checks at every handoff.
- Scale agent workers on queue depth, and track pass^k, not pass@k, for anything a customer touches.
The Unreliability Tax: Why Agentic Loops are Harder Than They Look
An agent is a model calling tools in a loop until the task is done. The model gets the attention, but the thing you ship is the harness: everything around that loop, from the prompt and the tools to the middleware that shapes what the model does. Moving from one request to a loop turns a stateless API call into a stateful run that has to survive tool failures, bad reasoning and its own bill.
Stevens calls the price of that the Unreliability Tax: the extra compute, latency and engineering you spend because a probabilistic model will sometimes fail. The bill has a shape, too. Each turn resends the whole history, so token use grows quadratically with the number of turns. By one estimate a single-agent loop costs roughly 4x the tokens of a plain chat exchange, and a UIUC study put multi-agent systems at 4 to 220 times the tokens of their single-agent counterparts.
None of that is a reason to avoid agents. It is a reason to build the harness with the same care as a payments service, because that is the kind of failure it has.
Anatomy of a Production-Grade Harness
The loop itself is short: assemble context, call the model, act on its decision, and go again until a stop condition ends the run. Keep its state in one typed object. LangChain's AgentState is a typed dictionary holding the conversation and any fields your tools and middleware need, and its message list is append-only: new messages are added, never replaced. Append-only history is what lets you replay a run and see exactly what the model saw.
Stop conditions are where most loops fail. A well-designed loop defines explicit exit criteria, and a production one needs three of them as hard guardrails: a maximum iteration count, no-progress detection, and a token or cost budget per run. The OpenAI Agents SDK defaults max_turns to 10, where a turn is one model call, and typical production values are 15 to 25. For no-progress detection, the cheapest signal is repetition: the same tool with identical arguments for a third consecutive turn means the agent is stuck, not working. The usual cause is a tool error the model never reasons about, so it tries the same call again.
import json
MAX_TURNS = 20 # production values are typically 15 to 25
MAX_TOKENS = 50_000 # per-run budget
REPEAT_LIMIT = 3 # same tool, same arguments, three turns running
def stuck_model(messages):
# Stands in for the LLM: it keeps asking for the same lookup.
return {"tool": "get_weather", "args": {"city": "London"}, "tokens": 1_200}
def run_tool(name, args):
return {"status": "error", "error_type": "upstream_timeout"}
def run_agent(task, model=stuck_model):
messages = [{"role": "user", "content": task}]
recent, spent = [], 0
for turn in range(1, MAX_TURNS + 1):
step = model(messages)
spent += step["tokens"]
if spent > MAX_TOKENS:
return {"stopped": "budget", "turn": turn}
if "tool" not in step:
return {"stopped": "done", "turn": turn, "answer": step["content"]}
call = (step["tool"], json.dumps(step["args"], sort_keys=True))
recent = (recent + [call])[-REPEAT_LIMIT:]
if len(recent) == REPEAT_LIMIT and len(set(recent)) == 1:
return {"stopped": "repeating", "turn": turn, "tool": call[0]}
result = run_tool(step["tool"], step["args"])
messages.append({"role": "assistant", "tool_call": step})
messages.append({"role": "tool", "content": json.dumps(result)})
return {"stopped": "max_turns", "turn": MAX_TURNS}
print(run_agent("What is the weather in London?"))Running it prints {'stopped': 'repeating', 'turn': 3, 'tool': 'get_weather'}. Without the repetition check, this model would burn all 20 turns on a dead upstream. Return the reason you stopped, not only the fact, because "repeating" and "budget" need different fixes.
Everything else hangs off the loop as middleware, one concern per piece. LangChain's human-in-the-loop middleware uses an after_model hook, which runs after the model responds but before any tool executes. That is the place to pause for a person before a destructive write or an expensive API call, without restructuring the agent.
Mitigating Failure: Hallucinations and Infinite Loops
The guardrails above end a loop that has stopped making progress. A loop that makes progress the wrong way is harder to spot, and retries make it worse. If the harness retries a failed call to a tool with side effects, it can send the message or take the payment twice. The fix is idempotency: give each tool call a stable key before it runs, so the harness can tell a retry from a new request.
import uuid
_results = {} # in production: a table with a unique constraint on the key
def charge(idempotency_key: str, amount_cents: int) -> dict:
if idempotency_key in _results:
return _results[idempotency_key] # a replay gets the first answer
if amount_cents <= 0:
return {"status": "error", "error_type": "invalid_input",
"message": "amount_cents must be positive"}
result = {"status": "ok", "charge_id": f"ch_{uuid.uuid4().hex[:8]}",
"amount_cents": amount_cents}
_results[idempotency_key] = result
return result
# The key comes from the task and the step, never from the attempt.
key = "task_123:step_4:charge"
first = charge(key, 5000)
retry = charge(key, 5000)
print(first == retry, len(_results))
print(charge("task_123:step_5:charge", -10)["error_type"])This prints True 1, then invalid_input. Two details matter. A replay returns the stored result rather than an error, so the model sees the same success both times and moves on. And the key is built from the task and the step, never the attempt, or every retry looks new. In production, make the check and the insert one atomic write, such as an insert against a unique constraint.
The same code shows the other habit worth copying: tools return structured errors instead of raising. A stack trace in the context tells the model nothing it can act on; invalid_input does.
In multi-agent chains, errors compound. Splunk summarises the Hallucination Snowball study, which injected 346 hallucinations into a four-agent financial pipeline: GPT-4o's detection rate fell from 72.0% at stage one to 50.9% at stage four, and another 23.7% reached the final output undetected. Boundary checks at each handoff reduced hallucination survival from 58.4% to 16.2%. A boundary check does not need to be clever. A schema check and a few domain assertions between stages catch most of what a final review misses.
Choosing an Orchestration Pattern: Cost vs. Accuracy
Start with one agent. Google Research's agent scaling study found that multi-agent coordination delivers +81% improvement on parallelizable tasks but causes up to 70% degradation on sequential ones. Most engineering work is sequential, and every extra agent pays for coordination in tokens.
When you do need several, the choice is measurable. A benchmark of four architectures on 10,000 SEC filings gives the trade-offs:
| Pattern | Use it when | What the benchmark found |
|---|---|---|
| Sequential pipeline | Steps depend on each other in a fixed order | The cost baseline |
| Parallel fan-out with merge | Parts of the work are independent | 1.84x lower latency than sequential, F1 up only 0.014 |
| Hierarchical supervisor-worker | You need oversight at a sane price | 98.5% of the reflexive F1 at 60.7% of its cost |
| Reflexive self-correcting loop | Accuracy outranks cost | Best field-level F1, 0.943, at 2.3x the sequential cost |
Hierarchical is the default I would pick: it sits on the cost-accuracy frontier across every model tested.
Control passes between agents through a handoff, and in OpenAI's pattern a handoff is just a tool. The function returns an Agent object instead of a string, and the loop switches to that agent. Because a handoff passes only what you choose to pass, write down four boundaries for every agent: the input it accepts, the decision it owns, the tools it may call, and the conditions that trigger a handoff. The boundary check goes on that handoff.
Infrastructure: The CPU-Bound Reality of the Loop
Inference runs on GPUs, but everything around it runs on CPUs: tool execution, context assembly, vector search, guardrails and orchestration. MCP adds to that, since each MCP tool call triggers data retrieval, transformation and formatting that runs entirely on CPU. Size the orchestration tier as its own service.
Observability via OpenTelemetry
You cannot debug a non-deterministic loop from logs. Trace it. Microsoft's Agent Framework emits an invoke_agent span as the top level span for each agent invocation and an execute_tool span for each function tool, with arguments and results as attributes. Follow the same shape in your own harness (pip install opentelemetry-sdk):
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter() # use an OTLP exporter in production
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("agent_harness")
def get_weather(city: str) -> str:
return "sunny, 22C"
def run_agent(task: str) -> str:
with tracer.start_as_current_span("invoke_agent weather_agent") as run:
run.set_attribute("agent.task", task)
with tracer.start_as_current_span("execute_tool get_weather") as tool:
tool.set_attribute("tool.args", '{"city": "London"}')
result = get_weather("London")
tool.set_attribute("tool.result", result)
return result
run_agent("What is the weather in London?")
spans = exporter.get_finished_spans()
names = {s.context.span_id: s.name for s in spans}
for s in spans:
print(f"{s.name} <- {names.get(s.parent.span_id) if s.parent else 'root'}")It prints the tool span as a child of the agent span. The in-memory exporter is there so you can see the structure; in production, swap in an OTLP exporter and the same spans reach your tracing backend. Traces of real runs are also where your evaluation tasks will come from.
Scaling and Error Handling
Do not autoscale agent workers on CPU utilization. It is a saturation metric: by the time utilization-based autoscaling reacts, users are already waiting. Queue depth is a leading indicator. It rises before latency degrades, because requests start queuing when every processing slot is busy, so new replicas start while the existing ones still respond normally.
Next Steps: Building a Reliable Evaluation Pipeline
Single-model benchmarks do not tell you whether a loop works. Amazon's agent teams evaluate task completion, reasoning, tool use and memory retrieval together, because that is where loops fail.
Track two numbers. pass@k measures the likelihood that an agent gets at least one correct solution in k attempts; it tells you what the agent can do. pass^k measures the probability that all k trials succeed, and Anthropic notes it especially matters for customer-facing agents, where users expect the same good result every time. At a 75% per-trial success rate, pass^3 is about 42%. High pass@k with low pass^k is a reliability problem, and the guardrails above are how you fix it.
Then build the suite:
- Start with 20 to 50 tasks drawn from real failures. You do not need a big dataset to begin.
- Grade code with code. Software is straightforward to evaluate: does it run, and do the tests pass?
- Simulate users. Amazon built an LLM simulator with virtual customer personas to cover the phrasings nobody writes by hand.
- Treat a 0% pass rate across many trials as a broken task or grader before you blame the model.
Run the suite on every prompt or tool change, and add each new production failure to it.
Sources
- Agents - Docs by LangChain, Docs by LangChain
- Human-in-the-loop - Docs by LangChain, Docs by LangChain
- The Agent Loop Decoded, blogs.oracle.com
- What Is the AI Agent Loop? The Core Architecture Behind Autonomous AI Systems, blogs.oracle.com
- The Hidden Economics of AI Agents: Managing Token Costs and Latency Trade-offs, Stevens Online
- The Anatomy of an Agent Loop, Steve Kinney
- Single-Agent vs Multi-Agent AI: When to Scale Your Dev Workflow, Augment Code
- OpenAI Swarm Multi-Agent Orchestration: The Complete Engineering Guide, Splunk
- Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies, arxiv.org
- Orchestrating Agents: Routines and Handoffs, OpenAI Developers
- CPU Inference and Orchestration - Amazon EKS, docs.aws.amazon.com
- Observability, learn.microsoft.com
- Demystifying evals for AI agents, anthropic.com
- Evaluating AI agents: Real-world lessons from building agentic systems at Amazon | Amazon Web Services, Amazon Web Services
aiarchitectureobservabilityengineering