Home / Blog /

Home / Blog /

Jev: the AI model that makes decisions

Jev: the AI model that makes decisions

Sep 19, 2026

Tools

Jev: the AI model that makes decisions

On 15 September 2026, a lab called TypeSafe came out of stealth with $40m in seed funding and a model named Jev. The headline capability is that it does not write.

The chat window was never the useful part

Every AI tool a business has adopted in the last three years has been built around conversation. That interface trained a generation of buyers to think of AI as a thing you talk to, and it quietly set the price and the speed of everything built on top of it. A model designed to produce paragraphs charges like one and makes you wait like one, even when all that was wanted was a yes or a no.

The expensive work in a small company is rarely the writing. It is the deciding. Is this lead worth calling back. Does this invoice look wrong. Should this enquiry go to the engineer, to accounts, or nowhere. Does this supplier email need someone today, or can it sit until Thursday.

None of those need a paragraph. They need an answer, quickly, consistently, several hundred times a week. Automating them has meant pointing an essay-writing machine at the problem and paying essay prices for a one-word reply.

What Jev does, and what it refuses to do



Jev does three things. It picks one option from a list defined in advance. It scores something against a scale defined in advance. It answers a yes-or-no question with a probability.

That is the whole product.

Because the possible answers are set by whoever is asking, the model cannot invent one. It cannot return a category that does not exist, a score outside the range, or a confidently worded sentence that turns out to be fiction. The failure that has made businesses nervous about AI is closed off by the design rather than patched over with a carefully worded prompt.

TypeSafe's chief executive, Diogo Almeida, is a former OpenAI researcher who worked on InstructGPT, ChatGPT and GPT-4, and co-invented RLHF, the training method that made conversational AI usable in the first place. The person who helped make AI talk has spent the last two years building one that will not.

One company launched it, the whole field is moving

Specialised models that do one narrow job have been quietly beating general-purpose ones at that job for a while, and 2026 is the year the gap became a purchasing decision rather than a research finding. For classification, extraction and routing, a small tuned model is routinely faster, cheaper and frequently more accurate than a frontier model. Sorting enquiries into four buckets does not get better because the model has also read the whole internet.

What follows is a tiered arrangement. Small specialised models handle the well-defined, high-volume work, and the expensive general model gets called only when something needs writing or reasoning through. Jev is the clearest published example of the first tier, because it has removed text generation entirely rather than merely discouraging it.

Six hundred milliseconds against a hundred seconds

TUSTRA runs an internal tool that assesses inbound opportunities against nine criteria. Until this week each of those judgements was made by a conventional conversational model. That layer was rebuilt on Jev and the same work run through both.

The old route took roughly a hundred seconds per assessment and cost about twenty-seven pence. The new one takes around six hundred milliseconds and costs a fraction of a penny. A full morning of testing came to less than a single penny in total.

The cost saving is pleasant. The speed changes the design of a system. At a hundred seconds and twenty-seven pence, you assess the things that obviously matter and let the rest through unexamined, because assessment is scarce and has to be rationed. At half a second and effectively nothing, it stops being scarce. Every enquiry, every invoice, every message can be scored continuously. That is a different kind of system rather than a cheaper version of the same one.

The number that mattered most was zero

On one test case, Jev returned a judgement with a confidence score of zero.

It had been asked to rate how commercially attractive a piece of work was. It looked at what it had been given, worked out that it genuinely could not tell, and said so, rather than producing a plausible number with nothing behind it. It was right. The question had been badly framed and the data did not contain enough to answer it.

This is the difference between a system that demos well and one that can run unsupervised. A system that quietly guesses is a liability, because wrong answers look exactly like right ones until something breaks in front of a customer. A system that raises its hand and says this one needs a person is something a real process can be built around.

Every judgement Jev returns carries that confidence figure, so the business sets the rule rather than the vendor. Above a threshold you choose, it acts on its own. Below it, a person looks.

The figures are the vendor's own, and the model is days old

Three things need stating plainly before anyone budgets for this.

The published performance claims have not been independently verified. TypeSafe's own testing reports the model running nearly 194 times faster and around 445 times cheaper than comparison language models, and cites a workflow cost of $0.39 against $3.31 for one competitor and $19.49 for another. Those are TypeSafe's numbers, published by TypeSafe, on the day it launched a product. The figures measured at TUSTRA are first-party too, taken on a single task over one morning, and a morning of testing is not a production track record.

It only does those three things. Anything requiring written output still needs a conventional model, so a business ends up running both. Jev decides, something else writes when writing is actually called for. Anyone selling this as a replacement for existing AI has misunderstood it.

The results depend on the quality of the question. The zero-confidence result above was a failure of framing, not of the model. Getting value out of this requires writing decisions down properly, which is a design skill rather than something that arrives free with a better model.

Where this earns its place

Lead qualification is the obvious one. Scoring every enquiry the moment it lands, rather than in a weekly review, means the good ones get a call back within the hour instead of on Friday afternoon when they have already booked someone else.

Triage and routing is the quiet one. Enquiries, emails and tickets going to the right person automatically, with genuinely ambiguous ones escalated rather than guessed at. Plenty of businesses have someone senior doing this by hand every morning without ever calling it a job.

Exception spotting is the one that pays for itself. Flagging the invoice, the order or the timesheet that does not look like the others, so somebody checks it before it becomes a write-off.

Verification is the one that compounds. Putting a decision model in front of other AI, checking its output before it reaches a customer and pulling in a person when confidence drops. That is how a system that currently needs watching becomes one that can be left alone.

What these have in common is that they are high-volume, low-glamour, and currently being done by people who should be doing something else.

The question to ask on Monday

The last three years taught everyone to evaluate AI by talking to it. If it holds a good conversation it must be clever, and if it is clever it must be useful. That was always a poor test. The conversation was the demo.

The better question is narrower and duller. Which decisions in this business are made hundreds of times a week, by a person, with a right answer that could be checked afterwards. Those are the ones where consistency matters more than eloquence and where speed compounds.

The useful question is no longer whether AI is clever enough to make the call. It is whether anyone has written down what the right call actually looks like.

Sources: TypeSafe AI launch announcement and company materials, 15 September 2026. Reporting by SiliconANGLE, AIwire and The AI Insider. Research on small and specialised models from NIST, BentoML and Turing Post. Every speed and cost figure attributed to TypeSafe originates with TypeSafe and has not been independently verified. The assessment timings and the zero-confidence result come from TUSTRA's own testing on 19 September 2026, measured on a single internal scoring task over one morning.

FREE EMAIL BRIEFING

The Tustra Briefing

AI in plain English for UK business.

One considered briefing covering what changed, why it matters, how businesses are using it, what to be cautious about and one practical action worth considering.

Important developments without daily noise

Practical UK business context

Honest case studies, security and adoption guidance

Free to join. Unsubscribe at any time. We will not sell your information.