Applied AI22/09/202617 min read

TypeSafe AI Jev: Structured Decisions in Milliseconds

We've been paying language models that generate entire paragraphs just to tell us "yes" or "no." We ask them to classify an email and get three lines of pleasantries before they reach the point. TypeSafe has just launched Jev, an AI model that does the exact opposite: it takes a state with bounded options, returns a structured decision with its probability distribution in milliseconds, and stops there. Making useful decisions that fast and that cheap, that's what matters to me. I tested it in two runs: 200 captures from real agency work and 81 additional requests covering marketing messages, operational urgencies, and evidence verification. Here's what I found, what works, and what doesn't.

Marketing Ultra Mascot

TL;DR: The no-nonsense summary

  • Jev doesn't write: it takes options and returns a decision (Choice, Score, or Noul). If you need text, you need a different model.
  • $0.042 per million input tokens: output is free. 200 requests using real captured material cost $0.0275 total.
  • Median of 0.39 to 0.77 seconds per decision: depends on the batch and option complexity.
  • 19 out of 20 agreements classifying synthetic agency messages, but stumbles with ambiguous inputs that lack context.
  • Guaranteed structure is not guaranteed accuracy: Jev always returns a valid category. Whether it gets it right depends on your criteria and your review.
In this article
  1. What Is TypeSafe AI's Jev
  2. How to Use TypeSafe AI's Jev
  3. What Happened When Testing Jev with 200 Real Captures
  4. Six Uses of Jev in Marketing (With Examples)
  5. How Jev Contributed to This Research
  6. Jev with GBrain and the Limits of AI Memory
  7. How Much Jev Costs and How Fast It Responds
  8. Jev vs. an LLM and the Rules That Always Apply
  9. Where Jev Falls Short
  10. Frequently Asked Questions About TypeSafe AI's Jev

What Is TypeSafe AI's Jev

Jev is an AI model from TypeSafe AI, announced on September 15, 2026. They call it a "System One Model": a model trained to return bounded decisions, not to generate free text.

What does that mean in practice? When you send it a question with fixed options, Jev picks one and tells you how confident it is in each alternative. It doesn't write. It doesn't explain. It doesn't add pleasantries.

It works with three question types, which TypeSafe calls primitives:

  • Choice: picks one option from several. "Is this message about sales, support, or billing?"
  • Score: applies an ordered scale. "On a scale of 1 to 5, how urgent is this incident?"
  • Noul: answers yes or no with a probability. "Does this excerpt support the claim in the report?"

In all three cases, the response is structured. ALWAYS. There's no risk of Jev returning malformed JSON, inventing a field, or adding an unsolicited paragraph of explanation. That's what TypeSafe sells as "no hallucinations": the structure is guaranteed.

But here's the catch: guaranteed structure doesn't mean the chosen category is correct. Those are very different things. Jev will return one of the options you gave it, yes. But whether that option is right for your business depends on how you defined the categories, the quality of the context you sent, and your own review.

How to Use TypeSafe AI's Jev

Access is through the TypeSafe API, with SDKs available in Python and TypeScript. It's also on Vercel AI Gateway, where in its first 24 hours it reached nearly 13% of paid teams and doubled the adoption rate of any previous gateway launch. That's a signal of initial interest, not a measurement of retention or global adoption.

Jev anatomy: Choice, Score, and Noul primitive channels with probability gauges converging to a guaranteed structured output.

The flow is always the same: give it context (the state), a question, and the possible options. Jev returns the chosen option and the probability distribution across all alternatives. No decoration.

The pattern: input, question, options, decision, action

Say a message hits your agency's shared inbox. The Jev cycle would be:

  1. Input: the message text as it arrives.
  2. Question: "What is the main intent of this message?"
  3. Options: sales, support, billing, cancellation, employment, clarify.
  4. Decision: Jev picks a category and returns its confidence in each option.
  5. Action: your system routes the message to the right team. Jev doesn't execute the action, it only decides.

That sequence takes less than half a second and costs a fraction of a cent. A generative LLM would do the same thing, but slower and more expensive, and it would throw in a courtesy summary nobody's going to read.

Note that these pieces are distinct: Jev is the model. The API is the service. The SDK is the library that simplifies the calls. Vercel AI Gateway is an intermediary that adds observability and caching. And a code agent skill is a set of instructions that can include Jev calls as part of a larger flow. They're not the same thing.

▶ Want to try it yourself?

Copy this and paste it into Claude Code, Cursor, or your favorite code assistant:

Install the TypeSafe AI SDK with pip install typesafe-ai. Set up the API key from https://typesafe.ai and test a Choice decision that classifies the message "I want to cancel my subscription" between support, cancellation, and billing. Follow the quickstart at https://docs.typesafe.ai/introduction/quickstart.

You don't need to know how to code. The assistant handles installation, configuration, and testing.

What Happened When Testing Jev with 200 Real Captures

The first test ran 200 requests on captures from real agency work. 40 for batch and category development, 160 for evaluation. Model pinned to jev-1.13.0. Sequential requests, no automatic retries.

The numbers:

  • 200 requests completed without API errors.
  • Median latency: 0.769 seconds per decision.
  • Estimated cost for the full batch: $0.0275. Less than three cents for 200 decisions.

On the 38 unambiguous provisional references I prepared before running, Jev matched all 38. I also cross-checked its decisions against historical Opus decisions on the same inputs. They agreed, which is not the same as saying Jev is better: I didn't re-run Opus under the same conditions or have comparable latency. An observation, not a controlled benchmark.

And the most useful thing I got from the experiment? Both models could reproduce questionable editorial rejections. That is: if the categories or criteria you defined are flawed, switching models doesn't fix the problem. It just executes it faster and cheaper.

After this first batch, I ran 81 additional requests (65 with public and synthetic data, 16 with evidence verifications) whose results I detail in the marketing and memory sections.

Six Uses of Jev in Marketing (With Examples)

Jev's promise fits any workflow where you need to classify, score, or make yes/no decisions many times, fast, without generating text. In agency marketing there are dozens of those decision points. I've worked through six in detail: three with my own testing on synthetic data and three as proposals with input/output design. As I cover in the guide to AI agents in marketing campaigns, a proposal that looks good and a flow that actually works are very different things.

1. Classifying agency requests (tested)

20 synthetic agency messages. Each arrives in the shared inbox: is it sales, support, billing, cancellation, employment, or does it need clarification? Jev matched the reference in 19 out of 20.

The one failure is the most interesting. Faced with the message "Budget" (a single word), Jev chose sales with 0.51 confidence. My reference expected "clarify," because a one-word message doesn't provide enough context. Is the sender asking for a new quote, checking on a pending one, or disputing a charge?

Who's right? That depends on how you design the categories. Jev isn't wrong to return sales, it makes the most probable decision with the information it has. The problem is that the information is insufficient. A well-built system would use that low confidence (0.51 versus the typical 0.87 on clear messages) to request more context before routing.

2. Operational urgency: anger is not an emergency (tested)

6 synthetic messages designed to separate legitimate frustration from a real emergency. An angry complaint because a report has the wrong colors. A store that can't process payments. A client threatening to leave over a billing error.

Result: 6 out of 6 agreements with the reference. Jev separated legitimate (but non-urgent) irritation from the kind of problem that shuts a business down.

These are six deliberately small, synthetic examples. They don't prove it works the same way in a real inbox with hundreds of variations. But they confirm that the categories and urgency criteria are well-designed for this type of distinction. The model responds to what you ask, if you ask poorly, you get a poor answer.

3. Research source curation (piloted)

While preparing this very report, I gathered social signals with Last 30 Days. Jev classified the resulting 39 excerpts into five types: reported result, demo/code, proposal, opinion, and insufficient information.

The output lives in a CSV with a link to each source and its confidence level. Useful as an auxiliary grouping tool. As a replacement for a researcher's judgment, not even close. Jev classifies the text it receives, it doesn't open links or verify original sources. More on the traps of this use case in the research section.

4. Reviewing claims in reports (tested)

8 claims about our own test results, checked against a pre-prepared evidence summary. Example: "Jev will cut any agency's hours in half." Jev rejected it as unproven.

It also rejected "perfectly identifies purchase intent in Spanish" and flagged as a contradiction the claim of having measured both models under equivalent conditions. That comparison wasn't done. 8 out of 8 matches with the provisional reference.

It's a small, curated batch, but it teaches something valuable: you can use Jev as a consistency filter between what a report claims and the evidence behind it. TypeSafe's official citation verification recipe goes deeper into this pattern.

5. Editorial overlap (proposed)

Before writing a new article: does one already exist that covers the same ground? The input would be the headline plus a summary of the new piece, alongside the headlines and summaries of candidate posts from the blog. Jev returns "new," "update," or "review."

I haven't tested it with real blog data, but the design fits as a preliminary step in any editorial workflow where the catalog grows faster than the team's memory. The important thing: you need to compare actual pages, not just headlines. A similar headline can cover a completely different angle.

6. Video clip pre-selection (proposed, with external reference)

Eric Siu presents a similar use case in his video on Jev for marketing work: evaluating transcript segments and deciding whether an idea is complete, a potential hook, or dependent on visual context. The input is the segmented transcript with episode context.

The limit is clear: selecting text is not the same as cutting a clip. Video editing needs image, audio, and pacing that Jev doesn't evaluate. But as a first filter to narrow an hour of footage down to the moments worth manual review, it makes sense.

The other 12 proposals (untested)

The full catalog covers 18 applications. The 12 not listed above are untested: they're input, decision, and limit designs that would need validation with real data before claiming anything about their usefulness.

ApplicationWhat Jev decidesWhat to watch for
Lead fitICP fit and review routingDon't use missing data as facts
Commercial intentBuy, compare, research, or supportAgreement with real sales data
Voice of the customerObjection, desire, friction, or alternativeConsistency across reviewers
Churn signalsOne-off issue or actual intent to leaveFalse churn alarms
Campaign commentsQuestion, complaint, purchase intent, or otherConversational context
SEO query intentIntent and need behind the queryDon't confuse label with volume
Brief complianceMeets requirements, missing info, or contradictsEvidence for each requirement
Ad inventoryOffer type, CTA, objection, and funnel stageDon't predict ROAS without data
Creative formatsCandidate format for the ideaBrand criteria outside the model
Creative feedback classificationCopy, image, offer, brand, or measurementMulti-dimensional comments
Internal task routingAssigned team or clarification neededAvailability rules
Content opportunitiesRecurring question or usable caseAuthor judgment before publishing

How Jev Contributed to This Research

To prepare this report I collected 49 URLs from different sources (GitHub, Hacker News, Reddit, YouTube). Jev classified the 39 excerpts with sufficient text into five categories: reported result, demo/code, proposal, opinion, and insufficient information.

The subsequent review showed why the text it receives isn't always enough. The Omi GitHub issue was classified as demo/code because it contains example code. But its actual adoption status is an open proposal, not a deployed integration. Something can have code and still be just an idea.

I also found that some search engine summaries mix user comments with the main thread body. Jev classifies what it receives, it doesn't open links or certify the original source. If the summary mixes things up, the classification inherits that mix.

This use served as an auxiliary grouping tool. Time savings weren't measured, and no source was automatically discarded based on what Jev said. The final call was human.

Jev with GBrain and the Limits of AI Memory

Where I've gotten the most value from Jev is evaluating fragments retrieved from an AI memory system: separating useful evidence from noise, flagging contradictions, and deciding whether a fragment is enough to answer a question or whether more searching is needed. TypeSafe's official RAG recipe describes exactly that pattern.

I tested it on 16 authorized cases: 8 retrieval (RAG) fragments and 8 claims checked against evidence. Model jev-1.13.0. Median: 0.385 seconds. Estimated cost: $0.000383 for the 16 cases. Agreement with provisional reference: 8/8 on RAG and 8/8 on claims.

But here's the caveat. Fragment R3 was real but incomplete: a slice of a note that didn't contain the necessary information. Fragment R-full was a manually prepared summary drawn from the complete report. These aren't two equivalent automatic retrievals. R3 arrived the way it would from a real search system; R-full is a hand-enriched condition.

If you work with semantic search or AI memory, engrave this: a downstream classifier doesn't fix an absence of evidence. If what you retrieve doesn't contain the answer, classifying it better won't invent the missing data. The system needs to be able to open the full note, retrieve another section, or reformulate the search. Filtering more doesn't mean knowing more.

Some correct decisions had moderate confidence (for example, 0.44 in one case). Don't turn those numbers into an automatic threshold chosen after seeing the results, that's overfitting dressed up as engineering.

How Much Jev Costs and How Fast It Responds

TypeSafe's published rate as of this report is $0.042 per million input tokens, with output free. These are conditions observed in the official launch post, not a guarantee of future pricing.

Mascot examines a compact precision sorter; a massive steam boiler performs the same sorting task at far greater scale behind him.

Here are the costs and latencies I measured during testing:

BatchRequestsMedianInput tokensEstimated cost
200 captures (first run)2000.769 sn/a$0.0275
65 public requests650.399 s36,163$0.00152
16 evidence verifications160.385 s9,109$0.00038
Second run total810.398 s45,272$0.00190

All requests were sequential, from a Python client over this machine's network. Cost is calculated from the token usage returned by TypeSafe's API, not from a reconciled invoice. And it doesn't include the actual cost of preparing the data, coding the tests, running the research, or reviewing the results.

A user on r/AI_Agents reports roughly one second per decision versus 4 to 14 seconds from an LLM with structured output. There's not enough detail to make it a reliable benchmark, but the order of magnitude is consistent with what I observed.

Jev vs. an LLM and the Rules That Always Apply

Since launch, the same question keeps coming up: "Does Jev replace ChatGPT?" No. They do different things, and the right choice depends on what you need:

NeedFirst option to evaluate
Compare numbers, sum costs, verify a literal stringDeterministic code
Apply a semantic taxonomy to many inputsJev (with evaluation and error review)
Write, explain, propose text, or reasonGenerative model (Claude, GPT, Gemini)
Search documents and retrieve sourcesSearch engine or retrieval system
Authorize actions or publishSystem rules and human judgment

The real value comes when you combine several pieces. TypeSafe's citation verification recipe is illustrative: the literal existence of a citation is checked with code (string comparison), but the semantic relationship between the citation and the claim is a task where Jev fits. You don't need to pay for an LLM call just to check string equality.

Syntax demonstrates this in their video: a chatbot that combines Jev's decisions with external tools and text templates. Jev picks the action; another component executes it. And Sam Witteveen explores how a small change in a support message's context can shift both the assigned category and urgency. The model responds to the text it receives, not to what the customer "meant."

What was NOT done in these tests: a controlled comparison against an LLM or an optimized rules-based classifier. Historical decisions and synthetic references don't substitute for one. To claim savings or superiority you'd need the same sample, human labels, per-class error metrics, and full-flow costs.

Where Jev Falls Short

281 requests and manual review. Here's what fails, or what launch enthusiasm has overstated:

Poorly defined criteria. If your categories are ambiguous or overlapping, Jev will pick one and move on. It won't ask "hey, are you sure these options make sense?" The "Budget" case illustrates this: the model chose sales because it's the most probable option given a single word. The problem wasn't the model, the input lacked sufficient context to decide well.

Confidence is not accuracy. A confidence score of 0.9 means Jev concentrates probability in that option, not that it's right 90% of the time in your use case. Noul, additionally, doesn't include the same confidence field as Choice and Score. Automating decisions with a threshold chosen after seeing a small batch is overfitting dressed up as engineering.

Missing evidence. Jev classifies what you send it. If what you send doesn't contain the necessary information, the classification is an educated guess. I saw this with partial RAG fragments and with search engine summaries that mixed body content with user comments.

It doesn't generate, explain, or search. This seems obvious, but it's worth repeating: Jev doesn't draft an article, write the email that goes to the client, query your CRM, or retrieve documents. If your workflow needs any of those things, you need another piece. Jev decides. Full stop.

What's missing. I haven't measured retention, I haven't run a controlled comparison against a rules-based classifier, I don't have full-flow production costs or human-labeled validation from a marketing team. The 18 application proposals are exactly that: proposals. Of the six I developed, three have synthetic or pre-prepared data with small batches, one is a grouping pilot, and two are proposals with no real data. None have production results.

Making useful decisions that fast and that cheap. That's what Jev promised and that's what I've seen in these bounded batches. It's not the future of marketing and it won't replace your team's judgment. It's a cheap, fast piece for a specific type of work. And what it does, you can measure. The interesting part starts when you stop asking what Jev does and start asking how many of your business's decisions you're overpaying for.


Frequently Asked Questions About TypeSafe AI's Jev

Is TypeSafe AI's Jev open source?

No. Jev is a proprietary model that runs on TypeSafe's servers through their API. The Python and TypeScript SDKs are public, but the training code, weights, and internal architecture are not available. There are community projects (such as hermes-jev-skills or typesafe-router) that integrate Jev into their flows, but these are not the model itself.

Do I need to know how to code to try Jev?

You need access to the TypeSafe API and an environment where you can run Python or TypeScript. If you use a code assistant like Claude Code or Cursor, you can ask it to install the SDK and run a test without writing any code yourself. The official quickstart guide covers the steps in under ten lines of code.

What does the confidence score Jev returns mean?

The confidence score in Choice and Score is calculated from the probability distribution across options: it indicates how much Jev concentrates probability on the chosen option, not the actual accuracy rate in your use case. Noul, the yes/no primitive, doesn't include the same field. Using these numbers as automatic thresholds requires evaluation with your own data, not extrapolation from someone else's batch.

Leave a comment

Your email will not be published. We review comments before showing them.