Charge the failed attempts and the work of repair before declaring a cheaper route.
A cheap model becomes expensive when somebody has to finish its answer. The first call can look like a routing win while the complete request accumulates another inference, repeated retrieval, and a review queue. A token-price comparison ends before those costs arrive. The useful unit is the work that passes an agreed acceptance test.
This week’s routing argument has moved from choosing an eligible destination to preserving progress during a handoff. The next question is whether the whole path earns its bill. Cost per accepted answer is total attributable operating cost for a defined request cohort divided by the number of answers that meet its requirements. Failed attempts belong in the numerator even though they contribute nothing to the denominator. With no accepted answers, report the spending and zero successes; the ratio is undefined.
Define acceptance before comparing routes. Consider a hypothetical documentation assistant that explains a software change from permitted source records. An answer must cite the applicable revision, support its explanation, preserve relevant qualifications, and arrive within the agreed deadline. A plausible paragraph or a successful HTTP response cannot satisfy that contract. If a person repairs the answer, record it as accepted after repair, with the repair charged to the route.
The denominator needs a boundary too. Count unique completed requests rather than every candidate response, retry, or intermediate artifact. Give a cohort enough time to finish, report unresolved requests separately, and retain abandoned work’s spending. Otherwise, one route can look efficient because its expensive cases are still waiting for review. Publish acceptance rate and coverage beside the ratio so a system cannot improve its apparent economics by dropping difficult work.
Here is constructed arithmetic, not a benchmark or a current price quote. Send one hundred requests through an economical model at one cent each. Forty require a stronger-model attempt costing another five cents each, and ninety requests eventually pass the fixed acceptance test. Model spending is three dollars, or about 3.33 cents per accepted answer. Sending all one hundred directly to the stronger model costs five dollars; if ninety-five pass, that is about 5.26 cents each.
Now assume the cascade also requires twenty one-minute reviews, while the direct route requires five. At an illustrative labor allocation of sixty cents per minute, those reviews add twelve dollars and three dollars respectively. Keep the same final acceptance counts, including any reviewer repairs: the cascade costs about 16.67 cents per accepted answer, against 8.42 cents for the direct route. The first model’s cheaper call survived the comparison. The route’s advantage did not.
Those assumptions cannot tell you which model to buy. They expose what the comparison has to include. Change first-pass acceptance, escalation frequency, reviewer effort, or the price of either call, and the result can reverse again. A review may be required by policy regardless of model quality; charge that common requirement consistently. Treat additional correction caused by a route as a separate burden rather than declaring every minute of review avoidable.
The surrounding software can move the bill even with the model held fixed. Liu and Han’s October 3 preprint, What Does a Harness Buy? Tokens, Mostly, reports cost per task differing by up to threefold across coding-agent harnesses under a common price ledger. The study pins Claude Code 2.1, mini-SWE-agent 2.4.6, and OpenCode 1.18 revisions, uses offline SWE-bench Verified tasks and a 300-call budget, and examines both locally served Qwen models and vendor API models. Its monetary analysis uses the API arms; local models contribute token counts. Repeated prompt and tool-schema input helps explain the spending differences. These are author-reported preprint results, not measurements reproduced here or proof about this documentation assistant.
Build the ledger around a request identifier that survives every attempt. Record the routing-policy version, attempted destination, reason code, model revision when available, and acceptance-rubric version. Attach retrieval, generation, validation, tool execution, and review as separate events with parentage. Each event needs its outcome, attributable cost, and elapsed time. The final destination should not inherit a clean bill merely because earlier attempts disappeared into another log.
Gateway accounting supplies part of that record. LiteLLM’s spend-tracking documentation describes calculated response costs and spend records, and recommends reconciling discrepancies through time ranges, token categories, and pricing data. That is useful input to the application ledger. It does not establish whether an answer passed, how long a reviewer spent correcting it, or what a paid retrieval service charged. Reconcile the model component with actual billing and keep unknown costs visible.
Retries need distinct failure labels. A transient transport failure, an invalid output, unsupported claims, and missing source evidence demand different remedies. LiteLLM’s fallback contract can move to another model group after configured retries. Application quality failures still need an explicit path. Estimate retry probability by failure class and attempt number; the next attempt is conditional on the previous failure, not another average request drawn from the whole workload.
Tool-call waste deserves its own field. If the economical route retrieves the same records three times, its lower generation bill has purchased extra database work and delay. Mark which retrieved artifacts remain usable by the next attempt. Charge repeated execution where it occurs and avoid charging a shared retrieval twice. A tool timeout also needs an uncertainty state: reconcile whether an operation completed before reissuing it, especially when another attempt could duplicate a side effect.
Cache reuse changes the path more substantially. LiteLLM distinguishes response caching, which can return a stored response without another model call, from provider prompt caching. The ledger should preserve that distinction, including lookup, storage, and validation costs. Reuse only when authorization, source revision, and answer requirements still agree. Report cold and warm traffic separately as well as the expected traffic mix; a benchmark full of repeats can flatter a route serving mostly new questions.
Elapsed time remains a separate constraint. Track time from request arrival to accepted delivery, with retrieval, inference, escalation, active review, and queue waiting visible. Waiting is not automatically billable reviewer time, and parallel event durations cannot simply be added to obtain wall-clock latency. Inspect the median and slow tail, then record how often the route misses its deadline. An answer that becomes correct after the user needed it has failed a time-bounded contract.
There is a limit to what this ratio can govern. An unsupported statement that slips past a weak check makes the route look both cheaper and more successful. Audit accepted outputs independently, distinguish costly error classes, and account for later retractions within a stated observation window. Keep authorization and high-impact error limits as hard constraints. A low average operating cost cannot purchase permission to cross them.
Compare policies on the same permitted workload with the same rubric, deadlines, and accounting boundaries. Preserve task slices such as exact lookups, conflicting evidence, missing records, and long context; changing their proportions can change the winner. Run paired comparisons where practical and repeat enough cases to expose variability. Include a capable single-model baseline and a simple rule before crediting a learned router for savings. A router recommendation alone cannot reveal the outcome of an alternative that was never executed.
I would start with one bounded workflow and a ledger that can explain every rejected or repaired answer. Separate marginal request spending from allocated infrastructure and maintenance, then show both when making an adoption decision. Replace the current route only when the candidate clears the same quality and containment requirements at a better complete operating cost. The cheapest call has earned nothing until the request is finished.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


