AI Decision Models: Knowing When to Bring in a Person
Customer support was the first place I wanted to use a large language model. Since the GPT-3.5 days, the difficult question has been knowing where its role should end.

Back in the GPT-3.5 days, customer support was the first use for a large language model that came to mind, and the one I most wanted to try in my own work. The hardest part to pin down was when a person should step in. How much of the service could AI reasonably handle, and where should its role end?
The recent wave of AI decision models brought that question back to me. These models specialize in judgments such as which team should receive a request, whether there is enough information to use a tool, or whether a case needs someone’s attention. Those are small steps in a conversation, but they can shape how long it takes to get help.
By the time someone types “I’d like to speak to a person,” they may already have checked the order, read the help page and explained the problem once. Another courteous reply is welcome. What they are hoping for, though, is some progress.
The person picking up the case may have a long chat history to read and an order record in a different system. It is easy for both sides to spend their patience on that gap. That is where I would like these models to be useful.
What AI decision models do inside a support flow
Consider a customer whose parcel is marked as delivered but has not arrived. The next step might be checking tracking details, asking for missing information or passing the case to someone authorized to deal with it. Before anyone writes another reply, there is a fairly specific choice to make.
A conversational model usually generates a text response. Decision models tend to return an allowed option, a score or a probability that software can use directly. Both can sit inside the same service. My article on Liquid AI d1 looks more closely at one example, including its input limits and local deployment requirements.
Classification and scoring have been part of machine learning for a long time. The development worth watching is how providers are packaging language understanding, task instructions and fixed-choice outputs into reusable services. That could mean fewer separate rules to maintain and less output to generate for a routine routing decision.
Several companies are betting on these small decisions
On October 9, 2026, TypeSafe announced an $870 million funding round at a $7.5 billion valuation, led by a16z. Its first model, Jev, returns structured judgments based on supplied material and questions. The size of that investment draws attention to AI decision models. I am interested in which everyday features might eventually come out of it.
Microsoft has also introduced Microsoft-Decision-1 for tasks including yes-or-no judgments, choosing between options and assigning ratings. Microsoft says Xbox Research used it to categorize more than 10,000 pieces of feedback and reviews. That is an understandable role for the technology: organize a large collection of opinions so people can find the issues worth exploring.
Cloudflare updated its Clef family on October 9, adding image, audio and video input with Clef-omni and lowering the price of Clef-flash. The products differ, but all are trying to become part of the repeated decisions inside software. People using the finished service may never see the model’s name.
Lower prices make small improvements easier to consider
The public prices are easier to compare once they use the same unit. As of October 10, 2026, Jev and Microsoft-Decision-1 both list input at $0.042 per million tokens. Clef-flash is $0.038. Tokens are the units used to meter model input; capabilities, context limits and the way inputs are tokenized still differ.

Here is a hypothetical cost calculation. Suppose each decision uses 500 input tokens in total, including instructions and answer options. A million calls would use 500 million tokens, costing $21 at the $0.042 rate. Jev and Microsoft’s model do not charge for output. That covers one decision step, with order lookups, subsequent conversation and staff time still part of the wider service.
In the DeepSeek and Huawei article, I wrote about the distance between a cheaper API and savings a customer actually notices. Here, I would like lower costs to make modest improvements worth building: noticing a missing detail earlier, for example, or getting a request to the right team before someone joins the wrong queue.
There are trade-offs. Cloudflare reduced the hosted Clef-flash context window from 64k to 24k tokens alongside the price cut. A short routing request may fit comfortably, while a long conversation could need a different arrangement. The useful comparison starts with the material the service actually handles.
A different task can change which model does best
Cloudflare’s own evaluation table offers a useful example. Jev leads the four listed models on When2Call, while Clef-omni has the highest BANKING77 score. When2Call examines decisions about when to use tools; BANKING77 concerns banking customer-service intent classification. The chart selects those two tasks to show the difference.

These are results reported by Cloudflare, and I have not independently reproduced them. The two tasks also use different metrics, so the comparisons belong within each panel. A model that sorts customer requests well may perform differently when deciding whether a tool call is appropriate. Even within one support flow, different steps may suit different models.
TypeSafe’s documentation says Jev currently performs best in English. That still leaves plenty to check in an English-speaking service: abbreviations, misspellings, specialist terms and messages with several requests at once. Mixed-language conversations add another variable. A small set of familiar cases would give a team a more useful starting point than assuming a benchmark will carry over unchanged.
The boundary I wanted to understand back then
Thinking back to that first idea for my support work, a refund is a useful way to explain the boundary. Recognizing a refund request, checking whether it meets the policy and actually issuing the money are separate steps. A model can help identify the request; the service still needs to define what it is allowed to do next.
A valid output can also be the wrong choice. A model might send a case to shipping even though the customer has already spoken to the carrier and now needs the seller’s help. Understanding that bit of history could do more good than producing the same category more quickly.
Some delays sit elsewhere: an integration is missing, an approval is required or the next person is busy. AI decision models can help route the work, but the rest of the process has to connect. If the result is one less page of history for a colleague to search through and one less order number for a customer to repeat, that is already a worthwhile improvement.
I would also like a request to speak to a person to be easy to honor. Someone should not have to work harder to persuade the software that they need help. Where the system is unsure, handing over what it already knows could give the next person a useful place to start.
That is what keeps me interested in AI decision models. I still like the idea of bringing AI into customer support, with clearer limits and better ways to pass the work along. A little less repetition could leave both the person asking and the person answering with more patience for the part that actually needs them.
Sources and prices checked on October 10, 2026. The cover is an AI-generated concept image. Charts reproduce the linked official data. The cost example is hypothetical, and benchmark results are vendor-reported; I have not run an independent model API evaluation.