"A model that never writes a sentence" sounded like a joke until we counted how many sentences our own stack generates for no reason at all. On September 15, TypeSafe AI introduced Jev — the first model in a new " System One is a new class of AI model (the first is TypeSafe AI's Jev) that returns a typed decision with a calibrated probability instead of text — no string generation or parsing, so it runs far faster and cheaper than a regular LLM. Full definition →" class: it takes a block of program state plus a set of questions with a known answer set, and returns not text but a typed decision with a calibrated probability, in 70–500 ms. Rather than read this as one more model launch, we walked our own five LangGraph is a developer framework for building AI agents that follow multi-step workflows with branching logic, rather than answering in a single pass. Full definition → flows — the same ones we render on our product pages — node by node and asked each decision point one question: is this a job for Jev, is it still work for an An LLM (Large Language Model) is the type of AI model — like the ones behind ChatGPT — trained on huge amounts of text to understand and generate human language. Full definition →, or is it not a model decision at all.
Why this article
We could have written "we tried the new model from TypeSafe" — there will be dozens of those after the launch, and none of them will have anything to do with our code. We'd rather answer the question we actually have: we run five production agents built on LangGraph, each with its own control-flow graph, and each graph has spots where the language model isn't generating content — it's picking one of a few known-in-advance options. That second category is exactly what Jev targets. So we went graph by graph, listed every such point, and gave it a verdict: candidate, stays with the LLM, or not a model decision at all. The full table is below, along with three reasons we aren't migrating anything yet.
What a System One model is, using Jev as the example
A regular LLM, even called as a classifier, does the same thing it does for a long piece of prose: it predicts A token is a fragment of text, usually a piece of a word, and it is the unit AI models measure length and charge by — bills are counted in tokens, not in questions. Full definition → by token, sequentially, until it has produced a string that then has to be parsed back into structure. Jev skips that step. It takes a block of state and a set of typed questions with a closed answer set, scores all of them in one parallel pass, and returns a value and a calibrated probability for each option — no string generation, nothing to parse on the way out. That's where the claimed 70–500 ms end-to-end latency and the 40–200x cost advantage over frontier models on this class of task come from.
The name isn't arbitrary. TypeSafe named the model after William Stanley Jevons, the economist behind the Jevons paradox: when the cost of something drops sharply, demand for it doesn't shrink proportionally — it explodes, because uses that made no sense at the old price suddenly do. TypeSafe's bet is that once a single decision costs a fraction of a cent and takes half a second, applications will start asking the model in places they don't bother today — not because an LLM couldn't handle it, but because at the old price and the old latency, nobody wanted to ask.
One thing is worth being precise about. "Can't hallucinate" in TypeSafe's own materials means "can't return a value outside the declared type", not "can't be wrong". That's a format guarantee, not an accuracy guarantee. On TypeSafe's own four-workflow benchmark, they report 67.8% accuracy — a number worth keeping right next to every "no A hallucination is an AI answer that sounds convincing but is simply untrue: the model fills a gap with a guess instead of admitting it does not know. Full definition →" claim, not instead of it.
Pricing, and what we can't compute with it yet
Jev costs $0.042 per million input tokens, and output is free — TypeSafe's own phrase is "too cheap to meter". That's a pricing shape none of the models in our own agent cost calculator has: the cheapest entry in our price table, gpt-5.4-nano, costs $0.20 per million input and $1.25 per million output, and the calculator's own data type assumes an output price always exists and is greater than zero. Concretely: computing the savings from switching to Jev inside our own tool would require changing the data structure that describes it first. That isn't an argument against migrating — it's an item on the list of things to do before anyone can put a number on the savings.
Our stack: five graphs, twelve decision points
On our product pages we publish LangGraph diagrams for five agents: the accounting assistant, the budget assistant, the e-marketing assistant, test Orchestration is coordinating several models, tools and agents so they work on one task in a defined order and hand results to each other. Full definition →, and the legal knowledge base. We went through each one and listed every node where the graph branches into a known-in-advance set of answers — not every processing step, only the points where someone (a model or plain code) picks one of a few options. That came to twelve such points across five graphs.
The split isn't even. Three points are generation or judgment of free-form content — Jev has nothing to offer there, since by definition it doesn't generate text. Six are closed-set classification or routing with a small option count — exactly the job Jev was built for. The remaining three aren't a model decision at all: user confirmation is human input, not inference; the personal-data filter runs fail-closed, which means it has to be deterministic by design; and one node the diagram simply doesn't say who decides today.
Verdict, node by node
| Product | Decision node | Answer options | Verdict |
|---|---|---|---|
| accounting-ai | Call a tool (every loop step) | 2 — yes / no | Candidate — highest frequency, since it fires every step, not every message |
| budget-assistant | Action type | 3 — answer / read / write | Candidate |
| budget-assistant | User confirmation | 2 — yes / no | Not applicable — this is human input, not a model decision |
| emarketing-ai | Which of 7 specialists takes the task | 8 — direct answer + 7 agents | Strong candidate — the highest per-message call volume in the whole stack |
| testing-ai | Which of 5 test agents to run | 5 | Unclear — the diagram does not say if this is an LLM call or trigger-based routing |
| testing-ai | Is the generated test valid | 2 — refine / done | Stays with the LLM — judging generated code |
| testing-ai | Is the generated checklist valid | 2 — refine / done | Stays with the LLM — judging generated content |
| legalka-kb | Message type | 3 — question / /suggest / KB review | Candidate, though today probably resolved without an LLM |
| legalka-kb | Does the knowledge base cover the question | 2 — answer / abstain | Strong candidate — today a cosine-similarity threshold, not a model |
| legalka-kb | Intent: question or edit | 2 | Candidate |
| legalka-kb | Is the revised legal text correct | 3 — invalid / ok / abort | Stays with the LLM — substantive judgment of text |
| legalka-kb | Personal-data filter (fail-closed) | 2 | Not moving — a safety gate has to be deterministic |
Three rows deserve more detail, because they show that "candidate" doesn't mean "any classifier will do".
Specialist dispatch in the e-marketing assistant is the most frequently called decision point in the whole stack — it fires on every user message, before any of the seven specialized agents (content, SEO, email, checklists, strategy, analytics, documents) gets involved. Today GPT-4o makes that call as part of a tool invocation. This is exactly the kind of work TypeSafe's own launch material describes: high volume, repeated decision, known answer set up front.
Which agent in test orchestration picks one of five specialized agents (test generator, bug detector, checklist generator, coverage advisor, flaky-test detector) based on what triggered the run — a code change or a PR. The diagram doesn't say whether that's an LLM call or plain routing on the trigger type, so we're treating it as a conditional candidate: if it's an LLM today, Jev fits; if it's routing on a PR label, there's nothing to replace.
Coverage in the legal knowledge base is the most interesting case in the whole audit, so it gets its own section.
The most interesting case: a threshold, not a model
Most "candidates" in that table are today an LLM call that a cheaper, faster call could replace. Coverage is different: judging by the diagram, the node sits right after "retrieve · cosine top-K" and asks whether coverage is good enough to answer, or whether the system should abstain and say "I don't know". That reads like a plain numeric threshold on a cosine-similarity score — hand-tuned, and probably adjusted more than once after the system either answered something it shouldn't have touched, or refused something it should have handled.
Replacing a hand-tuned threshold with a calibrated probability from a model that sees the whole question and the retrieved A chunk is one of the fragments a document is split into, so an AI can retrieve and quote the relevant piece instead of reading the whole file. Full definition →, not just one similarity number, isn't "same cost, faster" — it's a different decision mechanism, potentially more accurate than a threshold that knows nothing beyond vector distance. It's the one point in the whole audit where moving to a System One model would mean adding a layer of inference where none exists today, rather than swapping one model call for another.
What stays with the LLM, and why
Three points we're deliberately leaving alone: judging whether generated tests and checklists are valid (the "valid? refine" loops in the test generator and the checklist generator) and checking a revised piece of legal text in the knowledge base. All three formally have a closed answer set — two or three options — so at first glance they look just as fit for Jev as e-marketing dispatch. The difference is what the model has to evaluate to reach that answer: not system state described in fields and numbers, but freshly generated, free-form content — test code, a checklist, a paragraph of revised legal text. Judging whether that text is correct, semantically consistent, and doesn't introduce a new error requires reasoning over content, not classifying known-in-advance state. That's exactly the boundary TypeSafe states outright: Jev is for decisions over state, not for judging free-form content.
Three reasons we aren't migrating anything yet
Even for the strongest candidates — e-marketing dispatch and knowledge-base coverage — we have three reasons to wait, not three reasons to give up.
First, 67.8% accuracy on TypeSafe's own four-workflow benchmark is a number from their tests, not from our traffic, and on a seven-way routing decision a wrong call means invoking the wrong specialist, not just a worse answer. Before moving anything, we need our own test set built on our actual messages, not someone else's benchmark.
Second, the personal-data filter runs fail-closed — meaning it's supposed to block when in doubt, not let through. A safety gate of that kind needs deterministic, auditable behavior, not a calibrated probability that by definition is sometimes wrong. Even if that filter were LLM-based today — which the diagram doesn't confirm — it wouldn't be the first candidate for a probabilistic model.
Third, our own cost calculator has an EU data-residency flag, and Jev's launch material says nothing about processing in an EU region. For a customer in finance or legal, that isn't a technical footnote — it's a condition of even having the conversation.
What this article doesn't settle
To be honest: this audit is built from the diagrams we ourselves publish on our product pages, not from reading each agent's source code line by line. A diagram tells you a node exists and where its branches lead; it doesn't always tell you whether today's decision is made by a language model, a plain code condition, or a regex rule. Where a node sits directly downstream of a block explicitly described as a call to a specific model, assuming "it's an LLM" is reasonable. Where it doesn't — like the test-agent selection — we flagged it as uncertain instead of guessing in whichever direction made the table look tidier.
What we do first
Before moving any candidate, two things need to happen in order: extend the pricing structure in our own cost calculator to cover a model with free output — today it can't even be computed — and build a small test set on real messages hitting the e-marketing dispatcher, so we have our own accuracy number instead of someone else's benchmark. Only after those two steps does "how much would we save" become a question worth asking, because the honest answer today is: we don't know, and that's something to measure, not guess.
Summary
Jev lands exactly on the part of our stack we rarely write about: not text generation, but the quiet choices between a few known-in-advance options that an LLM handles today simply because we had no cheaper alternative. Of twelve decision points audited, six are real candidates, three stay with the LLM because they judge free-form content, and three aren't a model decision at all, or shouldn't be. The most interesting finding here isn't about cost or speed — it's that in at least one spot, the RAG (Retrieval-Augmented Generation) lets an AI look up your own documents before answering, so its replies are grounded in your data instead of guesswork. Full definition → coverage threshold, moving to a System One model would mean more than swapping an engine: it would mean adding inference where today there's only a hand-picked number.
Frequently asked questions
- How is Jev different from a regular LLM called as a classifier?
- A regular LLM, even asked for a single-word answer, generates it token by token, sequentially, and the result then has to be parsed. Jev scores an entire set of typed questions in one parallel pass and returns a value plus a calibrated probability directly, with no string generation or parsing — hence 70–500 ms instead of several to a dozen-plus seconds.
- If Jev "can't hallucinate", does that mean its answers are always correct?
- No. "Can't hallucinate" in TypeSafe's own materials means the model won't return a value outside the declared type — that's a format guarantee. On TypeSafe's own four-workflow benchmark they report 67.8% accuracy, so the model can be confidently wrong the same way a regular LLM can; it will just never return something that fails to parse into the expected type.
- Which node in our stack is the best candidate for Jev?
- Specialist dispatch in the e-marketing assistant — called on every user message, picking one of seven specialized agents. It's the highest call volume in the whole audit, with a closed, known-in-advance answer set — exactly the task profile TypeSafe designed Jev for.
- Why isn't the personal-data filter (pii_guard) a candidate if it's a simple yes/no decision?
- This node runs fail-closed — when in doubt, it's supposed to block, not let through. A safety gate like that needs deterministic, auditable behaviour, and a calibrated probability is by definition sometimes wrong. It isn't a question of how many options there are, but of what should happen when the model isn't sure.
- How much would moving the e-marketing dispatcher to Jev cost?
- We can't compute that today: our agent cost calculator assumes every model has a non-zero output price, and Jev doesn't have one. Before we can give a real number, we need to extend the calculator's pricing structure and build our own accuracy test set on real messages — only then would a comparison be more than a guess.