What Can We Now Afford to Notice?
Imagine three project teams, working separately, asking for help with the same requirement. One asks its AI assistant to explain the guidance. Another sends a question to a specialist. The third works around the uncertainty and asks someone to check afterwards. This is a hypothetical example, but consider what success would look like in each case. Everyone receives a useful answer, the work moves on, and each request can be considered resolved. The recurring difficulty may still go unnoticed.
Someone who saw all three requests might conclude that the teams needed more training. They might be right. But capable people could also be struggling to reconcile two versions of the instructions. In that case, more training would add work while leaving the contradiction in place. Bringing the requests together would give someone a chance to tell the difference, and to decide whether to help people apply the requirement or correct the information they had been given.
That possibility makes me interested in the cost of recognising similar difficulties across many requests. Language models can already classify text. The question is whether it is affordable to ask useful questions of each item as it passes through a service, often enough for the answers to change what happens next. A lower cost could make repeated interpretation worth trying in places where it previously seemed excessive.
Jev, from TypeSafe AI, is designed for specified questions with outputs that software can use directly. The allowed form of the answer is set in advance: a yes or no, a choice from a list, a score. It can make a judgement about language without writing an explanation in response. This makes it a possible component inside a service someone already uses. A constrained answer can still be wrong; the surrounding software simply knows what sort of answer it will receive.
There is an early practical example in Scour, a personalised content feed. In his 29 September account, developer Evan Schwartz reports asking 54 questions about roughly 1.1 million documents a month for less than $150 in model charges. He sends titles, URLs and summaries or opening snippets, rather than full documents. For his self-funded project, he says using even the cheapest language models for these checks would have been prohibitively expensive; Jev made them feasible. His questions include the expertise an article assumes and whether it is likely to remain worth reading, and he can add questions without retraining. He also spent much of a week experimenting, adjusting questions and checking answers. The model bill excludes that effort and the wider cost of running the application. His account gives a concrete reason to reconsider repeated classification, while leaving open what a large organisation would save or whether its service would improve.
Suppose we tried this in a shared support service, using the requests it already receives. Each request could be assessed against a few questions: does it concern this requirement? Is the person asking where to find guidance, or how to apply it? Does the person describe conflicting instructions? People may use quite different words for similar difficulties, which makes interpretation useful. That last question could flag a reported contradiction. Establishing whether two documents actually conflict would require examining the relevant versions; a description of someone’s difficulty cannot supply information they have left out.
The choice of questions matters before any results arrive. If we ask only which training someone needs, we have given the system no way to flag contradictory instructions. We would need questions that leave that possibility visible, and room for requests that fit none of our categories. Even then, the service would see only the requests that reached it. A colleague asked in passing, or a team that worked around the problem without asking, would be absent from the picture.
Software could group the answers and show that several teams had asked about applying the same requirement that week. Someone familiar with the work could inspect the original requests, check the guidance and investigate what the teams were trying to do. That is how a collection of classifications might become an understanding of a recurring problem. The grouping makes the investigation easier to direct; the model has not discovered the cause by assigning a category.
For the person asking, the same process could help them reach a relevant example or specialist sooner. Ordinary software can apply a clear rule, such as which support team covers a location. Jev could help interpret the substance of the request. The language model in an assistant could help explain the relevant guidance or prepare a question for a specialist. The person would notice whether the help fitted their problem and arrived sooner.
Across the service, inspection might reveal that an obsolete set of instructions remained easy to find. Correcting the guidance, withdrawing that version and putting a worked example where people encounter the requirement could remove the reason for future requests. If the instructions were sound but people struggled to apply them to unfamiliar cases, time with an experienced colleague might help. The useful response depends on understanding why the questions recur. Resolving each request individually would leave that question open.
I would begin with one recurring difficulty for which earlier recognition would allow someone to make a useful change. There needs to be time to inspect the results and act on them: sending every uncertainty to an already busy specialist could make the service worse. The cost of the trial would include checking, maintaining and responding to the system. An existing rule or a model trained for that narrow task may be sufficient, and requests that can wait could be processed together later. Jev gives a reason to reconsider the economics; the whole service still has to justify the effort.
In our hypothetical service, I would want to compare time to useful help and specialist effort with how things work today. I would also follow some cases through to the work people were trying to complete. A fall in repeat requests could mean the guidance had improved, but it could also mean people had given up asking. Checking what happened to the work would help distinguish those outcomes. These are questions for a trial, not benefits established by the Scour example.
The next team to encounter the requirement is where I would look for the result. Another good answer may be exactly what they need. If the earlier requests revealed contradictory instructions, though, I would want that team to find clear guidance and be able to get on with its work. The organisation would have used what it noticed to change the experience of someone who had not yet asked for help.