Skipping inference is useful only when the request still has an owner and a next state.
A router that can only choose a model has already decided to spend money answering. It can compare providers, capabilities, and latency while missing the cheaper decision: this request should not reach inference yet. The answer may already exist, the input may be incomplete, or nobody may have authority to act. A better destination list includes those conditions before model selection begins.
Yesterday’s routing argument put policy ahead of the provider dropdown. The next step is to give that policy somewhere concrete to send unfinished work. “No model” covers several different outcomes: complete the request through code, return an applicable cached result, ask for missing context, defer until a dependency recovers, deny a prohibited operation, or hand a decision to its owner. Collapsing them into one refusal makes the interface simpler by hiding the job that remains.
Consider a hypothetical release-preparation service. An engineer asks it to explain the changes in a build and prepare a deployment request. The service receives an authenticated identity, a project, a build identifier, and the engineer’s message. It can read permitted release records and produce a proposed explanation. Deployment approval belongs to a separate owner. This is a reference design, not a system I have benchmarked or deployed.
Start with what the service can establish without interpreting prose. Does the caller have access to that project? Does the build exist? Has an identical preparation request already completed? If access is denied, the protected operation stops. If identity cannot be verified because a dependency failed, retain an unresolved request rather than inventing an access verdict. Open Policy Agent’s default-rule documentation demonstrates an explicit default allow value of false. The application still has to distinguish a policy denial from its inability to obtain the evidence for a decision.
A completed result can avoid generation altogether. LiteLLM documents response caching that returns a stored response instead of calling the model again. That differs from retaining prompt context to make another inference cheaper. In this release service, reuse needs a stronger contract than matching the engineer’s wording: project and access scope, immutable build and source revisions, output requirements, and relevant policy version must agree. Current authorization must still pass before the cached explanation leaves storage.
Freshness follows the thing being answered. An explanation of an immutable build can remain applicable while an approval or environment status changes. Keep those records separate, with invalidation rules for each. A cached deployment approval cannot substitute for checking today’s permission and target state. Even an exact textual match is insufficient when the facts behind the text have changed. A semantic near match creates another uncertainty to evaluate, not a shortcut around those checks.
Some requests need no stored model output either. A build timestamp, artifact identifier, or list of changed files can come from trusted records through a maintained formatter. A model earns its call when the engineer needs interpretation, such as connecting several verified changes into a useful release explanation. The route should preserve that distinction. Do not invoke a language model to decorate an exact lookup merely because every other branch speaks in paragraphs.
Missing context creates a different destination. “Prepare the release for the payments service” is incomplete if two builds are eligible. Return the permitted candidates and ask which build the engineer means. An explicit field validator can trigger that question without inference. If language creates the ambiguity, a classifier or decision model may help identify it, but its own call belongs in the cost ledger. “No generative model” and “no inference” are different claims.
Clarification needs an exit condition. Persist the missing field, the original request, and enough authorized state to resume when the answer arrives. Validate the supplied build rather than treating any reply as permission to continue. If the user never responds, expire the preparation request visibly. Repeatedly asking an open-ended model whether it now understands can consume more work than a narrow form and still leave the required identifier unresolved.
A specialist judge can help with the remaining semantic branch. TypeSafe’s intent-routing example sends different classes to deterministic handlers, specialist LLMs, or humans. The useful architectural feature is that model generation is only one destination. Its example cutoffs are not measured operating thresholds for this release service. TypeSafe’s confidence definition describes distribution concentration; a decisive result cannot establish that the build evidence is complete or the caller can approve deployment.
Defer when the missing prerequisite belongs to the system. If the source-record service is unavailable, keep the request pending with a bounded retry budget, an expiry, and an owner after retries end. Tell the engineer what is pending and when to expect another state change. Sending the same incomplete packet to a more capable model cannot recover the unavailable records. A timeout should become an observable dependency failure, not a confident explanation of a release nobody inspected.
Human review is appropriate when the evidence exists but the decision requires judgment or authority the service lacks. A request to waive a deployment condition can go directly to the designated release owner. Send the exact build, relevant condition, proposed exception, and unresolved question. Track acknowledgement and elapsed wait, with a backup owner if the deadline passes. A human destination with no service expectation is an unmonitored backlog, and a growing backlog eventually becomes user harm.
Durable pause is available machinery. LangGraph’s interrupts documentation describes saving graph state through a checkpointer and waiting for external input before resuming. It also warns that the interrupted node restarts from its beginning, so preceding side effects can repeat. Put irreversible execution behind validated approval and protect it against duplicate execution. Persistence preserves the workflow’s place; it does not supply a reviewer, a response deadline, or permission to deploy.
Refusing is not automatically the safer outcome. A needless denial can block a legitimate release; an indefinite deferral can strand a correction the organization needs. Distinguish the harm of an unauthorized action from the harm of delayed permitted work. Hold the protected deployment while allowing authorized preparation or status retrieval to continue when those tasks are independent. An appeal path and a visible deadline belong beside the denial rule, especially when automated interpretation helped trigger it.
Evaluate this router with abstention included. Geifman and El-Yaniv’s 2017 selective-classification paper studies the tradeoff between prediction risk and coverage by allowing rejection. Its image-classification results do not certify a release router. The transferable question is how errors change as automatic coverage changes. Report errors among accepted cases, automatic coverage, unnecessary denials, clarification completion, deferred-request age, and human resolution time. A system can lower accepted-case error by sending nearly everything away while making the overall service worse.
The test fixtures should include a repeated request after access revocation, two valid builds, an unavailable source service, a stale approval, an exception nobody owns, and an ordinary exact lookup. Check the whole resulting state transition, including the reply and eventual resolution. Choose any semantic thresholds on development cases, then evaluate untouched cases. No universal cutoff or savings percentage follows from this design; reduced generation spends less only if routing, retries, and follow-up work do not consume the difference.
I would begin with explicit rules and a small set of named outcomes before adding a learned router. Keep completed-through-code and completed-through-cache distinct from waiting-for-context, waiting-for-dependency, waiting-for-owner, and denied. Record why each branch occurred and whether the request ultimately finished. That ledger makes it possible to compare against one approved capable model without rewarding a router for dropping the work it was supposed to handle.
The consequential decision is who owns the request after inference stops. A cache hit can close it. A clarification must reopen it when the missing field arrives. A deferral needs a clock, and a human handoff needs a person who accepts the work. Add those destinations before tuning which model wins. If the router chooses no model, it still owes the user a next state.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


