Jev is one of the most interesting AI launches this year. It’s also a good illustration of what financial services need from AI beyond speed and cost.
What Jev is
TypeSafe’s new “System One” is not a generative model.. You give it text and a set of questions with predefined answers. It returns an answer with a calibrated confidence score, in under a second and at a fraction of the cost of an LLM.
SuperBryn recently showed how even an LLM-as-a-judge can be broken down into simple classification questions that Jev answers on every call. Across 34 voice-agent metrics, 27 reached 96% accuracy or higher on calls held back for validation, and each judge cost a few cents to build. The seven that fell short had one thing in common: the answer depended on how different parts of the conversation connected.
What we’re seeing
To test Jev, we used it to create a customer vulnerability classifier – an essential component in a compliance-checking system. The appeal of Jev is that we can use vulnerability definitions directly as binary questions, and obtain responses with calibrated probabilities. Using a test set of synthetic conversations, manually annotated with vulnerability types, we showed that Jev performed well out-of-the-box. It was comparable with our own specialised BERT-based decision model, fine-tuned on a separate corpus of synthetic conversations.
But whilst Jev is built to answer any question about any text, our customers need something more specific: judgements about regulated conversations that they can stand behind. As we evaluate it, five things stand out.
What regulated firms need
Evidence is crucial. Jev returns an answer and a confidence score. In compliance, that’s not enough: our customers need to know why. Our FinLLM models are trained to give the answer and point to the evidence behind it, in the transcript or document, so a reviewer can check every finding for themselves. A decision model, by design, doesn’t do that.
Knowing what to ask is the hard part. Jev can tell you whether something was said but whether it meets a regulatory standard is a different question. Turning a requirement into precise, testable checks, and knowing what the answers mean under the rules, takes deep domain expertise. A risk warning being read out, for example, doesn’t tell you the customer understood it.
Financial services inputs are demanding. Our work involves long calls and long documents, where the answer is often spread across a whole conversation or a group of calls and documents. With a 32k-token limit and a literal reading of each question, this is where models like Jev may struggle. SuperBryn’s results point the same way: the metrics that fell short were the ones that depended on context across the call.
Transparency about the model itself. We don’t know how Jev was trained and for regulated firms and their risk teams, being able to account for the model behind a decision is as important as the decision itself.
Data residency. Jev isn’t yet hosted in the UK or EU. For regulated firms in the UK/EU, that’s not a detail. They need to know where their customers’ conversations are processed.
Where it fits
None of this makes the idea less interesting. Fast, low-cost decision models could have a real place in evaluation, and in classification tasks where speed and cost matter, such as checking whether a required disclosure was read.
So we’re doing two things:
• Continuing to evaluate Jev, to understand where this type of model works and where it doesn’t.
• Exploring whether our own in-house models of this kind, trained on financial services use cases and hosted by us, could complement FinLLM.
Fast, cheap decision models are probably here to stay. In financial services, the question is how to use them responsibly, and that starts with understanding the problem.