Choose the primitive by the job, and give uncertainty somewhere accountable to go.
A model should not get a vote on a rule your application already knows. If an account cannot access a resource, asking an LLM whether the request sounds reasonable adds uncertainty to a settled boundary. If a ticket needs a department, generating a paragraph before extracting a label adds machinery to a bounded choice. The first architecture decision is whether this step needs generation at all.
September asked how much useful intelligence a smaller machine could serve. October starts one step earlier: what kind of intelligence does the job require? A system that routes every request to its strongest model has skipped that question. It pays for interpretation even when the answer lives in a database, and it can mistake a persuasive explanation for evidence that an action belongs on the permitted path.
Consider a hypothetical support workflow for a software service. A customer writes that an export failed, mentions a recent plan change, and asks for help. The application needs to establish account access, identify the likely owner, gather relevant records, explain the failure, and decide whether an exception requires approval. Those are separate jobs. Giving one model the message and the tools does not make them one decision.
Start with the account boundary. Read the authenticated identity and current entitlement from trusted records. Compare the requested operation with explicit policy. Missing identity or stale entitlement should stop the protected action and trigger a refresh or clarification, rather than invite the model to infer permission from the customer’s tone. Open Policy Agent supplies a concrete example of policy decisions evaluated against structured input, with enforcement handled separately. A model’s recommendation should travel through that boundary.
The customer’s language is a different problem. “My export stopped after I changed plans” might belong to billing, technical support, or a queue that handles uncertain cases. A classifier can select among those destinations without writing a reply. If you have stable labels and representative examples, a conventional text classifier deserves a baseline before you buy a more general service. Scikit-learn documents text vectorization into features that such classifiers can consume. Their maintenance still includes labels, evaluation, and changing vocabulary.
Similarity search can narrow the possibilities when past incidents or help documents matter. A nearby document is evidence to inspect, not proof that this ticket has the same cause. Two requests can share vocabulary while demanding different actions. Keep the retrieval result distinct from the decision about what it supports, especially when the proposed next step changes account state.
Jev makes another option visible. TypeSafe describes it as a text-input decision model that returns typed answers and probabilities rather than generated prose. You supply the state and define the judgments. That contract fits questions such as which team should inspect the ticket, or whether the message reports an export failure. It does not produce the investigation or the customer explanation.
The interface matters because code can consume it directly. TypeSafe’s Choice primitive selects from supplied options, Score evaluates described levels, and Noul returns a yes probability. Put an uncertain or unsupported destination in the routing design when the named teams do not cover the input. A valid label makes the response easier to handle. It does not establish that the selected team is correct, or that the information supplied to the judge was sufficient.
The distinction becomes sharp at the policy edge. A semantic judge might recognize that a customer is asking for an exception. Code can compare the account’s dates and limits; a person can own the exception. TypeSafe’s documented Jev 1.13 limitations include numeric precision and date comparisons. Moving those checks into a specialized model would preserve the original mistake under a different endpoint.
Generation earns its place when the workflow needs an explanation, synthesis, or a proposed plan. A generative model can connect the customer’s description with retrieved documentation and produce a draft response. A harder case may need a stronger reasoning model to reconcile conflicting evidence or propose further diagnostics. Neither capability repairs missing records. If the export log is absent, the honest route may be to obtain it before diagnosing anything.
That is the decision ladder: known policy in code, bounded recognition in a classifier or decision model, interpretation and synthesis in a generative model, and consequential ambiguity with an accountable person. It is a way to assign work, not a mandatory sequence of network calls. A request can stop at the first check. Another can go directly to human review because no model has earned authority over the consequence.
There is a respectable argument for keeping the workflow inside one general model. Each added component creates an interface, an outage mode, and an evaluation burden. A low-volume prototype may cost more to maintain as a collection of classifiers and services than as one model with constrained output. The smaller primitive has to win at the system level. Replacing one call with several calls and a review queue is not automatically an improvement.
I would make that comparison against three tests: decision quality, decision cost, and failure containment. Quality begins with labeled cases that include ambiguous inputs and missing evidence, rather than a handful of clean demonstrations. For the support router, compare mistaken destinations, cases it declines to route, and the work needed to correct them. A routing mistake that delays a reply and a permission mistake that exposes an account require different acceptance criteria.
Probability needs its own evidence. TypeSafe defines Choice and Score confidence from the concentration of their answer distributions; Noul has no separate confidence field. Calibration asks a different question: across comparable predictions, do observed outcomes match the probabilities? Scikit-learn’s calibration documentation explains that comparison. A concentrated distribution can still favor the wrong answer. Set operating thresholds against held-out outcomes for the actual task, with separate treatment of costly error types.
Cost includes the whole route. Count retrieval, inference, retries, fallbacks, validation, and reviewer effort per accepted result. Record the elapsed time the customer experiences, including the wait after an uncertain case reaches a person. A cheap routing judgment that sends too many cases into investigation can increase the workflow’s bill. No provider price or isolated speed claim resolves that comparison without the workload.
Failure containment asks what a wrong answer can change. A mistaken department label might be reversible; a credential-bearing operation needs an independent check before execution. In the hypothetical support system, the judge proposes a destination, the application records the reason and evidence, and the receiving handler verifies its prerequisites. The service that applies an approved change holds the credential. A fluent response or a high probability cannot bypass that service’s policy.
Abstention has to reach somewhere useful. If records conflict, gather the missing evidence or place the case in a queue with a named owner and a response deadline. If the decision service times out, hold the protected action while retaining the request. Recheck mutable account state before any eventual execution. A review queue with nobody responsible merely relocates the failure and makes the dashboard look quieter.
Choose one bounded decision in a workflow you already understand. Write down its allowed outputs, trusted inputs, unacceptable errors, and uncertainty path. Compare the existing approach with a rule, a conventional classifier, or a specialized judge before adding autonomy. Keep the evidence that explains which cases each approach handles and which it leaves behind. The model deserves the work only after the decision has earned its place in the system.
Subscribe to Signal Over Noise for practical analysis of AI agents and automation: what to test, which constraints matter, and when a tool is worth using. Bring the next idea to one bounded workflow and check the result before expanding it.
If this was useful, forward it to one engineer who needs less noise in their feed.


