A harder request earns a different tool only when that tool can resolve the remaining uncertainty.
A difficult support request can need a database lookup, a classifier, and a person without ever needing a frontier model. The mistake is treating difficulty as one quantity. Missing permission, unfamiliar wording, conflicting logs, and an exception request are different problems. Spending more inference on all four makes the workflow expensive while leaving some of them unresolved.
Consider a hypothetical customer message: “My export failed after the plan change. Please reset the limit and explain what happened.” The service must establish access, identify the responsible queue, retrieve evidence, explain the failure, and handle a requested account change. Each step needs its own output and acceptance condition. The decision ladder assigns those jobs to the smallest adequate primitive; it does not send every request through every rung.
Begin with the branch that admits no probabilistic override. Authenticated identity, current entitlement, exact quota arithmetic, and permitted operations come from trusted records and policy. Open Policy Agent illustrates the separation between evaluating policy against structured input and enforcing the resulting decision. In this design, the account service enforces the boundary. Missing or stale records trigger a refresh or a hold, never a model’s guess about what the customer deserves.
Rules have a small computational bill when the inputs are already available, although database reads and policy distribution still add latency. Their evidence is boundary tests, explicit treatment of missing fields, and review by the policy owner. Their maintenance burden is keeping definitions and records synchronized. The characteristic failure is an obsolete or incomplete rule that produces a perfectly repeatable wrong answer. Determinism makes behavior inspectable; it does not make policy correct.
Language begins at the next branch. If the only job is finding a documented error identifier, a lexical matcher may suffice. Keep a quoted identifier from the customer distinct from the same identifier in a trusted log. If the job is assigning varied descriptions to stable queues, compare a supervised text classifier with that simple baseline. Scikit-learn documents TF-IDF text features and logistic regression classification, an ordinary combination worth testing before adding a general model.
The classifier earns automatic routing with representative held-out tickets, per-queue error analysis, and unfamiliar-topic cases. A compact local classifier can avoid a hosted round trip, but labeling, retraining, and changing vocabulary are recurring costs. An unseen category may still receive a familiar label. Give unsupported inputs a review path, and test the threshold for accepting a route rather than assuming the highest score is adequate. The router selects an owner; it does not diagnose the export.
Evidence retrieval is a separate job. Embeddings can find incident notes whose wording differs from the ticket, while a cross-encoder can score the query and each candidate together. Sentence Transformers documents this retrieve-and-rerank arrangement: broad retrieval narrows the set before more expensive pair scoring. Indexed embeddings shift some work to document ingestion; reranking adds work per candidate. Neither component creates a missing log or establishes that a similar incident caused this failure.
Measure retrieval against known relevant documents and realistic near misses. Maintain the index when documentation, access controls, or embedding models change; filter evidence to the authenticated account’s permitted scope. A reranker cannot recover a document the retriever omitted. In the example, an incident about a different plan might rank highly while being inapplicable. Check product version, timestamps, and observed failure conditions before passing a passage to the next stage.
A specialized decision model belongs where the remaining output is bounded but its meaning requires language. Jev could select technical support, billing, or manual triage from defined criteria, or judge whether the message requests an exception. TypeSafe’s Choice, Score, and Noul primitives return option distributions, positions on described levels, and yes probabilities. Independent questions share state; application code combines the results. This is an alternative to the classifier when the decision contract benefits from language-defined criteria, not a mandatory call after classification.
Its acceptance evidence is a labeled workload, threshold and abstention analysis, and tests of misleading wording. Maintenance includes the question definitions and model version: TypeSafe currently maps its aliases to Jev 1.13, with version pinning available. Hosted requests add network, timeout, and retry costs; a smaller output contract does not establish an end-to-end speed win. The vendor’s jaggedness guide documents numeric and date weaknesses and susceptibility to adversarial state. Keep quota arithmetic and authorization outside the judge, even when its answer looks decisive.
Generation enters when the service needs a new explanation. A small generative model can draft a response from the verified export event, applicable documentation, and proposed next step. An enumerated route can also accompany that draft when combining the jobs proves simpler. OpenAI’s Structured Outputs documentation distinguishes schema constraints from ordinary JSON validity and warns that structured responses can still contain mistakes. A valid field containing an invented cause fails the task.
Test the small model on factual support, complete qualifications, refusal to invent absent evidence, and response usefulness. Its cost depends on input and output length, concurrency, and whether serving is local or hosted. Maintaining it means preserving prompts, model revisions, retrieval contracts, and semantic checks alongside schema validation. The common failure is an explanation that sounds finished before the investigation is finished. The customer reply stays a draft until those checks pass.
Frontier reasoning earns the next branch only when additional interpretation could settle the case. Conflicting logs or several plausible mechanisms may justify a stronger model proposing diagnostics. Compare accepted resolutions against the small-model baseline, including fabricated explanations and unnecessary tool proposals. Budget the complete attempt, including reasoning, retrieval, retries, and human correction; there is no universal latency ranking implied here. More capability brings another prompt, version, and tool-contract surface to maintain. If the necessary event record does not exist, fetch it or stop instead of buying a longer speculation.
Human review can enter from any branch. In this example, a requested limit reset outside ordinary entitlement goes directly to an authorized account owner, even if every model agrees. Send the original request, relevant records, proposed change, policy boundary, and unresolved question in one evidence packet. Give the queue a response deadline, a backup owner, and an appeal path. Human capacity has a labor bill and a waiting-time distribution; overload and inconsistent decisions are failure modes worth measuring, not proof that the queue should be automated away.
The concrete flow is therefore access check, accepted route, permitted evidence retrieval, checked explanation, and separately authorized action. A known error can use a maintained reply template and skip generation. An unsupported topic goes to triage; insufficient logs trigger an evidence request; a policy exception goes to its owner. A timeout preserves the ticket and holds protected changes. The account service rechecks mutable state immediately before applying an approved operation, then records the result.
Thresholds connect the components without pretending their scores mean the same thing. A retrieval score, classifier probability, and Jev distribution cannot share one generic confidence cutoff. Scikit-learn’s calibration guide describes comparing predicted probabilities with observed frequencies. Choose acceptance boundaries on separate validation cases, accounting for error costs and reviewer capacity, then report performance on a held-out set. Until that evidence exists, semantic routes remain proposals. No threshold in this reference architecture has been measured here.
There is a credible objection: a low-volume service may be easier to own with one constrained generative call than with several specialized dependencies. The ladder allows that outcome. Compare total cost per accepted ticket, elapsed time including review, missed costly errors, and the work of maintaining the route. Evaluate the simplest credible baseline first. Replace it only where another primitive handles a documented failure or reduces the complete operating bill without weakening containment.
Take one actual workflow and write down the next decision, its allowed outputs, trusted evidence, expensive mistake, and owner when it cannot proceed. Start with the lowest rung that can meet those requirements, then keep the cases where it fails. Those cases tell you whether to improve the evidence, change the primitive, or hand the decision to someone with authority. The next model call should have a job description before it has a budget.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


