A cheaper route earns traffic only after quality, deadlines, and failure boundaries survive review.
A router can save money and still fail its promotion review. The rejected candidate might answer routine questions well, finish most requests quickly, and produce an attractive average bill. None of that excuses sending a restricted document to an unapproved endpoint. A scorecard earns its place when it keeps that failure visible instead of averaging it away.
The routing argument needs a decision at the end: give this policy traffic, revise it, or keep the existing route. Quality, cost, latency, and failure belong on the same review record, but they cannot all be traded at the same exchange rate. I would use separate acceptance gates before comparing the candidates that survive. The cheapest eligible route can win; an ineligible route cannot buy its way back with savings.
Consider a hypothetical assistant that answers engineers’ questions from internal manuals and incident records. Routine lookups, conflicting evidence, and restricted documents share an interface but have different acceptance conditions. Its scorecard compares three policies: one approved capable model, a simple rule that separates known lookups from difficult interpretation, and a learned router. These are proposed comparison arms, not systems benchmarked here. Each receives the same permitted evidence and is judged against the same task contract.
Start the record with its identity. Name the workload, document revision, permitted destinations, model versions, routing configuration, evaluation window, and accountable owner. Keep the cases used to choose thresholds separate from those used to report the result. A scorecard without that context is hard to reproduce and easy to reuse after its assumptions expire. A change to the retrieval source can invalidate a result even when the router’s configuration stays unchanged.
Quality starts at the delivered answer. A routine lookup must cite the applicable manual and preserve its qualifications. An incident explanation must separate an observed event from a proposed cause. A missing record should produce an evidence request or an owned hold, rather than a plausible explanation. Matching a reference model choice is a useful routing diagnostic, but it does not establish that the resulting answer passed those requirements.
Report counts by task slice alongside the overall result. Keep a representative traffic set for the expected operating mix and a separate challenge set for conflicting revisions, incomplete records, and prohibited destinations. The challenge set tests specific boundaries; its proportions do not estimate their frequency in traffic. Record accepted coverage, rejected answers, and independently discovered errors among accepted answers. Otherwise, a router that declines most hard work can make its quality number look excellent while moving the burden to somebody else.
RouteLLM’s project guidance recommends calibrating on queries resembling the incoming workload and warns that the realized model split changes with the query distribution. Its evaluation framework also supplies public benchmark comparisons. Those are useful starting points, not acceptance evidence for this assistant. A desired share of strong-model calls is a spending control; the scorecard must still show which requests became acceptable and which errors survived.
The cost gate uses the complete policy’s bill. Include routing, retrieval, generation, retries, fallback attempts, and active human correction within a stated accounting boundary. Preserve the spending on rejected work in the numerator when calculating cost per accepted result. Show the number of accepted results beside that ratio, and leave it undefined when none passed. Yesterday’s accounting argument becomes useful here because the review must decide whether savings survive the same quality gate.
Keep request spending separate from allocated infrastructure and maintenance. A low-volume workflow may spend more maintaining a learned router than it saves in inference. An organization might still choose it for a documented quality advantage, but that is a different justification. Record unknown charges and the billing reconciliation window. A precise-looking total built from incomplete records should remain an incomplete cost result.
Latency needs its own gate because a correct answer can arrive after it is useful. Measure from request arrival to accepted delivery, including retrieval, validation, retries, and any review queue. Show the distribution for completed requests and count deadline misses and unfinished requests separately. Do not let requests that never finished disappear from the timing record. Parallel work contributes to elapsed time through its actual execution path, not the sum of every component duration.
Google’s SRE monitoring guidance explains why averages can conceal slow tails and why latency deserves attention even when a request eventually succeeds. In this assistant, a slow incident explanation might still be usable, while a late operational lookup might miss its purpose. Report timing by those task slices and by first-choice versus repaired completion. Choose the deadline from the workflow’s need before examining which candidate looks fastest.
Failure containment is the gate most likely to vanish behind a green success rate. Exercise a provider timeout, an exhausted retry budget, stale permissions, missing evidence, and an unavailable approved destination in a controlled test environment. These are proposed tests, not tests performed for this article. For each fixture, state the allowed result before execution and retain the observed route, held work, and recovery owner afterward. A fixture that merely produces some final answer has not proved containment.
LiteLLM documents retries followed by fallback to another configured model group. That mechanism gives the test something concrete to exercise; it does not certify the replacement’s eligibility for the document. The assistant should keep restricted material within the approved destination set through every attempt. If that set becomes unavailable, the expected outcome may be an owned hold with a deadline. Successful failover is a failure when it crosses the boundary the original route was required to preserve.
Enforcement needs evidence beyond a scorecard row. Open Policy Agent separates policy decision-making from enforcement: software supplies structured input and consumes the decision. The application must actually constrain dispatch. Inspect that boundary and test its treatment of absent or stale inputs. An evaluation with no observed disclosure supplies evidence about those cases, but it is not proof that a learned router can replace the enforcement mechanism.
The review can now make a narrower decision than declaring one router best. Promote the routine lookup slice if it meets the declared quality, cost, deadline, and containment requirements. Keep disputed incident explanations on the existing route if the evidence remains weak. Hold promotion when the sample is too small, a costly error lacks coverage, or accounting is unresolved. The decision record should name the permitted slice, remaining exclusions, reviewer, and conditions that trigger another review.
There is a reasonable objection to four gates: measurement and review have costs too. A small service with a satisfactory single-model route may gain little from this comparison. Start with the few failure cases and constraints that would actually change the owner’s decision, then add detail where uncertainty remains. The scorecard is a record of evidence, not a demand to build an observability platform before answering a question. It should make a simpler baseline easier to defend when complexity earns nothing.
After promotion, retain the cases that justified it and a route back to the baseline. Revisit the decision when a model alias moves, permissions change, traffic shifts, or repeated repair starts consuming the savings. Name who can stop the candidate and what evidence they need. A promotion review is useful only if its failure conditions remain actionable after the meeting ends. Give the router traffic it has earned, and keep the authority to take that traffic back.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


