Cost per Successful Task: how to measure what an AI agent really costs

Token prices say what a call costs, not what a finished job costs. How to calculate Cost per Successful Task from production logs, and which levers lower it.

TL;DR

  • Cost per Successful Task (CST) is everything you spent, failed runs included, divided by the tasks that passed your acceptance check.
  • Raise the success rate before you chase cheaper tokens: it moves CST further.

Token pricing hides the real expense of autonomous agents

A price sheet tells you what a call costs. Your budget cares what a finished job costs, and for agents the two drift apart three ways.

First, tokens are not the whole bill. Uvik's framework comparison puts LLM API calls at 40-60% of total agent operating cost in most production deployments.

Second, the scaffold around the model changes both the spend and the result. The same Claude Opus 4 scores 64.9% inside HAL's Generalist Agent scaffolding and 57.6% inside HuggingFace's Open Deep Research framework. On simple single-tool-call workflows, CrewAI carries up to 3× the token footprint of LangGraph or LangChain.

Third, failures compound. Chrono Innovation's arithmetic: an agent at 95% accuracy per step finishes a 10-step workflow correctly roughly 60% of the time. You pay for the other 40% and get nothing back.

If a run costs $0.40 and 60% of runs succeed, each success costs $0.67. The token dashboard shows the first figure; finance sees the second. WhatLLM's leaderboard says as much: WhatLLM notes that its cost per benchmark task "is not the measured cost of successfully completing your workflow".

Define Cost per Successful Task with the CLEAR dimensions

CST = total spend over a period, failed and retried runs included ÷ tasks that passed the acceptance check in that period.

Define "passed" first: the check a human reviewer would sign off on, automated at the end of the task. A task that fails it counts in the numerator only.

The CLEAR framework (Cost, Latency, Efficacy, Assurance, Reliability, from a November 2025 paper Uvik summarises) is a checklist for what CST must capture:

CLEAR dimensionWhere it lands in CST
CostNumerator: tokens, tools, hosting, framework fees
LatencyNumerator only if you price it, for example compute time or a person waiting
EfficacyDenominator: the success rate
AssuranceMeasure success in production. The paper finds a 37% average gap between lab benchmark scores and production performance
ReliabilityNumerator: every retry is paid for

The same paper finds cost varies up to 50× across agents achieving similar accuracy levels. Accuracy alone will not pick your agent; CST will.

Calculate CST from your production logs

  1. Tag every call with a task ID: LLM calls, tool calls and retries.
  2. Price each task: tokens at list price, tool fees, and its share of hosting.
  3. Record one pass or fail per task from the acceptance check.
  4. Divide over a fixed window, say a week, so one bad batch cannot swing it.

WhatLLM's advice holds for cost as well as time: "Measure the full workflow, including tool latency, retries, and failures; token generation speed alone does not predict completion time."

from dataclasses import dataclass
 
@dataclass
class TaskRun:
    llm_cost: float    # tokens x list price, all calls and retries
    tool_cost: float   # external APIs, search, sandboxes
    infra_cost: float  # this task's share of hosting and orchestration
    succeeded: bool    # passed the acceptance check
 
def run_cost(r: TaskRun) -> float:
    return r.llm_cost + r.tool_cost + r.infra_cost
 
def cost_per_successful_task(runs: list[TaskRun]) -> float:
    successes = sum(r.succeeded for r in runs)
    if successes == 0:
        raise ValueError("no successful tasks in this window; CST is undefined")
    return sum(run_cost(r) for r in runs) / successes
 
# One week: 6 passed; 4 failed, their costs including retries
runs = [TaskRun(0.30, 0.05, 0.02, True)] * 6 + [TaskRun(0.45, 0.08, 0.02, False)] * 4
print(f"Cost per run: ${sum(run_cost(r) for r in runs) / len(runs):.3f}")
print(f"CST:          ${cost_per_successful_task(runs):.3f}")
Cost per run: $0.442
CST:          $0.737

Failed runs cost more because they retried. Per run, $0.44; per finished task, $0.74.

Lower CST: raise the success rate first, then cut spend

Fix the denominator first. At the same spend, lifting success from 60% to 80% cuts CST by a quarter. A 25% cheaper model, with tokens at half the bill, saves about 12%. Chrono Innovation's pattern: "Define workflow phases with explicit checkpoints. At each checkpoint, a validation step confirms the accumulated state is coherent before the next phase begins." A run that fails the first checkpoint stops spending there.

Pick the framework on measured CST. Framework overhead varies by workload, so measure it on your own tasks. Uvik reports teams getting 40-50% LLM call savings on repeat workflows through stateful caching patterns. Work that does not repeat will not see it.

Parallelise independent subtasks. For 50 documents at 3 seconds each, sequential processing takes 150 seconds; parallel processing takes 3 seconds plus aggregation overhead. That lowers CST only if you priced latency; otherwise it buys speed, not savings.

Choose single-agent or multi-agent by measurement. My rule: start with one agent, and add more only when all three hold:

  • the subtasks are independent enough to run in parallel,
  • better prompts and checkpoints no longer raise the single agent's success rate,
  • a trial of the multi-agent version shows lower CST on a sample of the same tasks.

Cap the spend per task. Set a token budget and a retry limit, and alert when weekly CST rises above baseline. A run stopped by the cap counts as a failure.

Start this week: add task IDs to your logs, write the acceptance check, and compute last week's CST. That is your baseline: a change that does not lower it does not ship.

Sources

aifinopsagenticcostengineering

All writing