Every destination needs an eligibility rule, an acceptance test, and an owned way out.
A model router should be explainable without asking a model to explain itself. An operator needs to know which evidence made a destination eligible, why the request went there, and what happened when its answer failed. A fluent rationale cannot substitute for those records. Four explicit tiers can make the system inspectable, provided the tiers describe operating contracts rather than a ladder every request must climb.
Consider a hypothetical assistant that answers engineers’ questions from internal manuals and incident records. Some requests need an exact field, some need a short explanation, and some involve conflicting evidence or a decision the assistant cannot own. The proposed router has four destinations: deterministic code or an applicable cache; a small local or specialized decision model; an economical general model; and a frontier model or accountable human. This is a reference architecture, not a deployment I have measured.
The request enters through an authenticated service, not directly through the router model. That service resolves the caller’s access, identifies the task, retrieves permitted source records, and constructs a bounded evidence packet. Include immutable source revisions, sensitivity labels, required output shape, deadline, and acceptance-rubric version. Keep absent fields explicit. If the question concerns two possible incidents, the system should ask which one before spending inference on an ambiguous packet.
Apply eligibility before ranking destinations. A privacy rule might permit only local processing for one document class, while another permits specific contracted endpoints. A model score cannot add an endpoint to that set. Open Policy Agent’s documentation separates policy decisions from enforcement: software supplies structured input and consumes the result. In this design, dispatch enforces that result, including during fallback. Missing authorization produces an owned hold or denial, not a more expensive interpretation.
Tier one handles what trusted code can finish. An incident timestamp comes from a permitted record through a formatter. A cached explanation can return only when caller scope, source revisions, output requirements, and relevant policy still agree. Recheck authorization at delivery; yesterday’s permission is not part of an immutable answer. Emit EXACT_FIELD_LOOKUP or CACHE_SCOPE_AND_REVISION_MATCH from the rule that actually passed. An invalid cache entry records CACHE_INVALIDATED and continues with the current evidence, rather than quietly supplying stale prose.
Tier two handles a bounded uncertainty that remains after those checks. A small local model could map the engineer’s wording to a supported task class or extract an incident identifier. A specialized decision model could choose among named destinations without drafting the answer. These are distinct jobs: selecting an explanation route does not itself complete an explanation request. Preserve the intermediate result, validate identifiers against accessible records, and emit BOUNDED_TASK_CLASS only when the configured operating boundary passes. Unrecognized language or conflicting selections produce CLASSIFIER_ABSTAIN and an explicit next state.
The threshold belongs to this workload. RouteLLM’s threshold guidance recommends using queries resembling incoming traffic and warns that the realized model split changes with the query distribution. That supports workload-specific evaluation, not a universal confidence cutoff. Choose tier-two boundaries against labeled development cases, then test untouched cases and costly errors separately. A well-formed probability or a valid extraction is evidence about the model’s output contract, not proof that the request has been understood correctly.
Tier three receives tasks requiring ordinary synthesis from sufficient evidence. Suppose an engineer asks how a documented configuration change affects one component. The economical general model receives the relevant approved excerpts and a fixed output contract, not the entire incident archive. Emit SUPPORTED_SYNTHESIS_WITHIN_CONTRACT when task shape, evidence sufficiency, context size, and permitted endpoint match. “Economical” is an evaluation result to establish for that workload, not a promise attached permanently to a model name or advertised token price.
Context locality constrains that choice. Keep retrieval and source ownership in the application, and pass references to retained artifacts between attempts rather than asking each model to rediscover them. A reference does not make evidence available to a remote model: dispatch must resolve and transmit the approved subset when needed. A summary may omit the qualification that determines the answer, so retain its source mapping and make original permitted excerpts available to the acceptance process. Carry forward useful work without treating an earlier model’s interpretation as a new trusted source.
A small model is not automatically a private model. “Local” needs an actual execution boundary; a hosted specialist may send the packet outside it. Before remote dispatch, select only necessary evidence, apply the approved disclosure policy, and validate the resulting packet against the question’s requirements. Redaction that removes the decisive fact creates REMOTE_PACKET_INSUFFICIENT. That branch holds for authorized review or uses an eligible local route. It does not send the original material to the frontier endpoint because the sanitized version was inconvenient.
Tier four separates difficult interpretation from authority. Conflicting but complete incident evidence might justify CONFLICTING_EVIDENCE_REQUIRES_FRONTIER if a permitted capable model has earned that task slice. Missing records still require records, not frontier reasoning. A request to approve an operational exception produces OWNER_DECISION_REQUIRED and goes to a named reviewer with the evidence, unresolved question, deadline, and permitted responses. Record acknowledgement and expiry. A waiting human queue is unresolved work, and a stronger model cannot become the absent owner.
Acceptance sits after every completed candidate, including a cache return. For this assistant, require the applicable revision, supported factual claims, preserved qualifications, correct access scope, and timely delivery. Deterministic checks can verify identifiers and citation targets; semantic support needs an appropriate evaluator and independent sampling. An answer passing a schema check has not passed the entire rubric. If evidence is absent, record SOURCE_EVIDENCE_MISSING and request it. If interpretation fails with adequate evidence, record SEMANTIC_ACCEPTANCE_FAILED before considering an eligible stronger route.
The recovery budget covers the request, not each tier independently. For a first implementation, I would permit at most one additional generation attempt after an initial candidate fails, with the actual timeout and cost ceiling chosen by the workflow owner. This is a proposed policy, not a measured optimum. A transient read failure can use bounded backoff; an unsupported answer needs a different remedy. LiteLLM documents retries followed by fallback to another model group, but the application must preserve eligibility and acceptance across that transition.
Start another attempt only when the remaining deadline can accommodate execution and validation. Coordinate recovery in one layer instead of letting the client, gateway, and application each spend a fresh budget. Amazon’s retry guidance warns that layered retries multiply load and that a timeout does not establish that side effects failed. This assistant is read-only; if tools later change state, reconcile ambiguous outcomes or use a documented idempotency contract before repeating them. Budget exhaustion emits RECOVERY_BUDGET_EXHAUSTED with an owner and next state.
Build the audit event from those actual checks. Record request identity, policy and model versions, protected evidence references, evaluated routing features, selected branch, acceptance outcome, attempt parentage, and attributable cost and elapsed time. Reason codes come from application-controlled transitions, not a generated story about hidden deliberation. Keep rejected candidates and fallback history. Limit access and retention for the protected references; an audit log should not become a second, less controlled incident archive.
The tiers need not execute in sequence. A known owner decision should bypass generation. A task already evaluated as difficult can go directly to an eligible frontier model. Test the complete transitions against a capable single-model baseline and a simple rule, including stale permissions, revoked cache access, missing evidence, classifier abstention, timeout, and an unacknowledged review request. Promote only the slices that clear the same acceptance and containment gates. The router earns its complexity when every delivered answer and every stopped request has a traceable reason and an accountable owner.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


