Prepared by SwarmSystemThe Jev Playbook · As of 2026-09-20 · Jev 1.13
A playbook for a new kind of model

Jev: fast answers, and how sure

This page is for a developer or a technical operator on a small team. You pay a monthly bill for model calls. Somewhere in your code you make the same small judgment thousands of times: is this urgent, which bucket does this go in, does this match. This playbook shows what Jev does, whether it fits that job, what it costs, where it breaks, and how to start without getting burned.

What it is

A model that returns numbers, not words. You send text and a typed question. It sends back a probability and a confidence. It never writes a sentence.

Who it is for

Anyone with a decision code can act on. Routing, sorting, flagging, ranking, checking. Not writing, not summarizing, not explaining.

What it costs

$0.042 per million input tokens. Output is free. A 300 token message costs about a thousandth of a cent. Answers come back in 70 to 500 milliseconds.

What breaks

Counting, math, dates, tricky wording, hostile text. The vendor lists nine failure modes on its own site. They are all on this page.

Written from public sources before the author held an API key. Jev is in early access. Everything here is dated and linked so you can check it.

Is this for you

Decision model or a language model, and when

Most teams reach for a language model for every AI job. A language model writes. That is slow and pricey when all you needed was a yes, a no, or a pick from a list. Here is the split.

Use a decision model when
  • The answer is one of a list you already know: billing, technical, or sales. Yes or no. Low, medium, or high.
  • Code acts on the answer. It routes, sorts, flags, ranks, or gates. No human reads prose.
  • You need it in under a second, thousands of times.
  • You would like to know how sure the model is, so you can send the shaky ones to a person.
Use a language model when
  • You need words: a reply, a summary, a draft, a rewrite.
  • You need a reason: an explanation a customer, an auditor, or a manager will read.
  • The answer is not on a list you could write down in advance.
  • The job needs several steps of thinking before the answer.

Plenty of systems need both. The common shape: a decision model sorts and gates at the front, a language model writes only for the cases that need words, and a person gets the ones nobody is sure about.

What it returns

See one answer and you understand all of it

This is TypeSafe's own quickstart example, unchanged. You send a state (any text or JSON) and a set of questions. Every question has a type.

POST https://api.typesafe.ai/v1/systemone
{
  "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
  "model": "jev-latest",
  "questions": {
    "department": { "type": "choice", "instructions": "Which team should handle this",
      "criteria": { "billing": "Payment or subscription issues", "technical": "Bugs or integration problems", "sales": "Pricing or account questions" } },
    "frustration": { "type": "score", "instructions": "How frustrated the customer appears",
      "criteria": [ "Calm, just stating facts", "Frustrated but civil", "Very angry, strong language" ] },
    "is_urgent":  { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity" }
  }
}

Response
{
  "model": "jev-1.13.0",
  "answers": {
    "department":  { "type": "choice", "choice": "technical", "confidence": 0.78,
                     "probabilities": { "technical": 0.85, "billing": 0.15, "sales": 0.0 } },
    "frustration": { "type": "score", "score": 1.0, "confidence": 1.0 },
    "is_urgent":   { "type": "noul", "noul": 0.95 }
  }
}
Noul

A yes or no question

Returns the probability that the answer is yes, from 0 to 1. There is no separate confidence. A number near 0.5 means the model finds yes and no about equally likely.

Choice

Pick one from a list

Returns the pick, a probability for every option, and a confidence. You can list up to 255 options. Add a "none of these" option when nothing might fit, or it will pick anyway.

Score

Rate against described levels

Returns a weighted spot on the levels you describe. Good for ranking and for "is it over the line." Not good for measuring: a 1.4 is not "40 percent of the way to angry."

All the questions in one call run at the same time, so asking ten costs about the same time as asking one. The vendor calls this fan out. Question ids are for your code; the model never sees them, so put the whole meaning in the question.

Try it from a terminal

Get a key from the TypeSafe console, export it, and run the vendor's own curl. If you do not have a key yet, the console has a playground that takes the same state and questions.

curl -X POST https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
  {
    "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
    "model": "jev-latest",
    "questions": {
      "urgency": { "type": "noul", "instructions": "Does this message express urgency?" }
    }
  }
EOF

Pin a version in production. jev-latest points at jev-1.13.0 today and will move. A threshold you tuned on one version is a guess on the next. There are also official Python and JavaScript SDKs, and a Claude Code plugin that teaches a coding agent the API.

Confidence

Use the number to decide who decides what

The probability tells you what the model thinks. The confidence tells you whether to act on it. TypeSafe's guidance splits it into three bands. Move the slider to see how a threshold changes who handles the work on a sample of 100 decisions.

0.90
0.50
Code acts0
Review0
A person0

Sample of 100 made up confidence values. The shape is illustrative; your real spread will differ. The point is the trade: a higher act line means fewer mistakes acted on and more work for people.

BandWhat to doExample
High (above 0.9 for anything with real stakes)Let code act. Log it.Route the ticket. File the invoice. Archive the newsletter.
MiddleAct with a check, or flag it for a quick look.Show the suggested bucket, let a person confirm with one click.
Low (below 0.5)Do not act. Send it to a person or ask for more information.Put it in the queue with the model's guess visible.

The vendor's own words: "Start with conservative thresholds, test with your own data, and adjust as you observe results." Two more rules from their docs. Different actions in the same system should have different lines, because the cost of being wrong differs. And confidence tells you how concentrated the probability is, not how right the model is. A wrong answer can be confident.

Where it breaks

Nine ways it can go wrong, in plain English

TypeSafe published a page listing where Jev 1.13 fails and what to do instead. That page is the most useful thing on their site. Here it is in plain language, with their fixes.

#It struggles withWhat that looks likeDo this instead
1Reading between the linesIt answers the question you wrote, not the one you meant. "Not unhelpful" trips it.Write the exact condition. Put edge cases in the criteria.
2Math and countingIt cannot count items reliably or tell if two numbers are close.Do the math in code. Ask one yes or no per item, then add them up yourself.
3Dates"Which date is earlier" and "is this inside the window" are unreliable.Ask it to pull out the month, day, and year as a pick from a list. Compare in code.
4Multi hop questionsDouble negatives and "the owner of the account that sent this" cost accuracy.Point straight at the field. One hop.
5Long, noisy inputAccuracy drops as unrelated text piles up.Trim in code first. Send only what the question needs.
6Hostile textText written to steer the model can move the answer.Precise criteria, test the ugly cases. Never use it as your only security check.
7Mixed signalsWhen the question and the criteria disagree, it gets confused.Make them say the same thing, plainly.
8Consistency across question typesThe same question asked as a yes or no and as a two option pick gave different numbers. Two opposite yes or no questions did not add to 1.Do not move a threshold from one question type to another. Word each question to mean exactly what you want.
9Writing anythingIt cannot draft, summarize, or explain.Use a language model.

So keep it out of

How to start

Five steps, and the model earns every one

Code owns the decision. The model supplies a number. A number earns trust through a record.

1

Pick one decision

One place where code or a person makes the same small call many times a day. Sorting support mail. Flagging urgent messages. Ranking search results. Not the scariest one. The one with the most volume and the least damage when wrong.

2

Write the question like a contract

State the exact condition. Give every option a one line description. Add a "none of these" or "not stated" option. Send only the fields the question needs. Keep it in one file with a version number, because a threshold tuned on one wording is a guess on another.

3

Run it in shadow

Call it on real traffic and record the answer, the probability, and what actually happened. Change nothing. Do this until you have at least 100 examples with a human label. At the published price this costs less than a cup of coffee.

4

Compare it to what you do today

Agreement with your labels. Whether the 0.9s were right about 90 percent of the time. Cost per thousand decisions next to today's cost. If it is worse than the rule it would replace, stop. The record is still useful.

5

Promote one rung at a time

First show the number beside the human's decision. Then, after two clean weeks and a written sign off, let code act above a line you chose from your own data. Keep one switch that turns it off. Keep the record. You can replay it against any other model later.

Unproven

What has not been proven yet, as of this week

Early access

You may be on a waitlist

Keys come from the vendor console. The API is hosted only. No weights are published and there is no self hosting.

The headline numbers

193.6x faster, 444.6x cheaper

Measured by the vendor on workflows the vendor wrote, against reference answers that were the average of two large language models. The vendor calls these "the higher end of real world gains." Independent tests show 5 to 20x on speed and cost, with slightly lower accuracy.

Calibration

No independent curves yet

The whole value is that 0.9 means right about 90 percent of the time. TypeSafe trains for that. Nobody outside has published the curve. Build your own from your shadow record.

Price

It may be subsidized

The vendor says the price may be subsidized and expects it to fall, and cannot prove otherwise. Design so you can leave: keep the record, keep the interface, pin the version.

Accuracy

One independent number to hold

On an invoice processing test, Jev scored 61.8 percent against 79.1 for a frontier language model. Faster and cheaper is real. Smarter is not the claim.

Data

Your text leaves your machine

The vendor says it does not train on customer requests. Zero retention is an enterprise feature. Send the fields the question needs, not the whole record.

Sources

Check every claim on this page for yourself

The vendor

Press and independent tests

If you are not a developer

Someone will pitch you AI that decides

What happens when it is wrong, and who sees it? If the answer is "it is rarely wrong," walk away. If the answer is "a person sees every call it is not sure about, and here is the record," you are talking to someone who has read this page.

SELF CHECK: 0 TOKENS · 0 EM DASHES · 0 CONSOLE ERRORS · UNVERIFIED
Book a call