This page is for a developer or a technical operator on a small team. You pay a monthly bill for model calls. Somewhere in your code you make the same small judgment thousands of times: is this urgent, which bucket does this go in, does this match. This playbook shows what Jev does, whether it fits that job, what it costs, where it breaks, and how to start without getting burned.
A model that returns numbers, not words. You send text and a typed question. It sends back a probability and a confidence. It never writes a sentence.
Anyone with a decision code can act on. Routing, sorting, flagging, ranking, checking. Not writing, not summarizing, not explaining.
$0.042 per million input tokens. Output is free. A 300 token message costs about a thousandth of a cent. Answers come back in 70 to 500 milliseconds.
Counting, math, dates, tricky wording, hostile text. The vendor lists nine failure modes on its own site. They are all on this page.
Written from public sources before the author held an API key. Jev is in early access. Everything here is dated and linked so you can check it.
Most teams reach for a language model for every AI job. A language model writes. That is slow and pricey when all you needed was a yes, a no, or a pick from a list. Here is the split.
Plenty of systems need both. The common shape: a decision model sorts and gates at the front, a language model writes only for the cases that need words, and a person gets the ones nobody is sure about.
This is TypeSafe's own quickstart example, unchanged. You send a state (any text or JSON) and a set of questions. Every question has a type.
POST https://api.typesafe.ai/v1/systemone { "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.", "model": "jev-latest", "questions": { "department": { "type": "choice", "instructions": "Which team should handle this", "criteria": { "billing": "Payment or subscription issues", "technical": "Bugs or integration problems", "sales": "Pricing or account questions" } }, "frustration": { "type": "score", "instructions": "How frustrated the customer appears", "criteria": [ "Calm, just stating facts", "Frustrated but civil", "Very angry, strong language" ] }, "is_urgent": { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity" } } } Response { "model": "jev-1.13.0", "answers": { "department": { "type": "choice", "choice": "technical", "confidence": 0.78, "probabilities": { "technical": 0.85, "billing": 0.15, "sales": 0.0 } }, "frustration": { "type": "score", "score": 1.0, "confidence": 1.0 }, "is_urgent": { "type": "noul", "noul": 0.95 } } }
Returns the probability that the answer is yes, from 0 to 1. There is no separate confidence. A number near 0.5 means the model finds yes and no about equally likely.
Returns the pick, a probability for every option, and a confidence. You can list up to 255 options. Add a "none of these" option when nothing might fit, or it will pick anyway.
Returns a weighted spot on the levels you describe. Good for ranking and for "is it over the line." Not good for measuring: a 1.4 is not "40 percent of the way to angry."
All the questions in one call run at the same time, so asking ten costs about the same time as asking one. The vendor calls this fan out. Question ids are for your code; the model never sees them, so put the whole meaning in the question.
Get a key from the TypeSafe console, export it, and run the vendor's own curl. If you do not have a key yet, the console has a playground that takes the same state and questions.
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"urgency": { "type": "noul", "instructions": "Does this message express urgency?" }
}
}
EOF
Pin a version in production. jev-latest points at jev-1.13.0 today and will move. A threshold you tuned on one version is a guess on the next. There are also official Python and JavaScript SDKs, and a Claude Code plugin that teaches a coding agent the API.
The probability tells you what the model thinks. The confidence tells you whether to act on it. TypeSafe's guidance splits it into three bands. Move the slider to see how a threshold changes who handles the work on a sample of 100 decisions.
Sample of 100 made up confidence values. The shape is illustrative; your real spread will differ. The point is the trade: a higher act line means fewer mistakes acted on and more work for people.
| Band | What to do | Example |
|---|---|---|
| High (above 0.9 for anything with real stakes) | Let code act. Log it. | Route the ticket. File the invoice. Archive the newsletter. |
| Middle | Act with a check, or flag it for a quick look. | Show the suggested bucket, let a person confirm with one click. |
| Low (below 0.5) | Do not act. Send it to a person or ask for more information. | Put it in the queue with the model's guess visible. |
The vendor's own words: "Start with conservative thresholds, test with your own data, and adjust as you observe results." Two more rules from their docs. Different actions in the same system should have different lines, because the cost of being wrong differs. And confidence tells you how concentrated the probability is, not how right the model is. A wrong answer can be confident.
TypeSafe published a page listing where Jev 1.13 fails and what to do instead. That page is the most useful thing on their site. Here it is in plain language, with their fixes.
| # | It struggles with | What that looks like | Do this instead |
|---|---|---|---|
| 1 | Reading between the lines | It answers the question you wrote, not the one you meant. "Not unhelpful" trips it. | Write the exact condition. Put edge cases in the criteria. |
| 2 | Math and counting | It cannot count items reliably or tell if two numbers are close. | Do the math in code. Ask one yes or no per item, then add them up yourself. |
| 3 | Dates | "Which date is earlier" and "is this inside the window" are unreliable. | Ask it to pull out the month, day, and year as a pick from a list. Compare in code. |
| 4 | Multi hop questions | Double negatives and "the owner of the account that sent this" cost accuracy. | Point straight at the field. One hop. |
| 5 | Long, noisy input | Accuracy drops as unrelated text piles up. | Trim in code first. Send only what the question needs. |
| 6 | Hostile text | Text written to steer the model can move the answer. | Precise criteria, test the ugly cases. Never use it as your only security check. |
| 7 | Mixed signals | When the question and the criteria disagree, it gets confused. | Make them say the same thing, plainly. |
| 8 | Consistency across question types | The same question asked as a yes or no and as a two option pick gave different numbers. Two opposite yes or no questions did not add to 1. | Do not move a threshold from one question type to another. Word each question to mean exactly what you want. |
| 9 | Writing anything | It cannot draft, summarize, or explain. | Use a language model. |
Code owns the decision. The model supplies a number. A number earns trust through a record.
One place where code or a person makes the same small call many times a day. Sorting support mail. Flagging urgent messages. Ranking search results. Not the scariest one. The one with the most volume and the least damage when wrong.
State the exact condition. Give every option a one line description. Add a "none of these" or "not stated" option. Send only the fields the question needs. Keep it in one file with a version number, because a threshold tuned on one wording is a guess on another.
Call it on real traffic and record the answer, the probability, and what actually happened. Change nothing. Do this until you have at least 100 examples with a human label. At the published price this costs less than a cup of coffee.
Agreement with your labels. Whether the 0.9s were right about 90 percent of the time. Cost per thousand decisions next to today's cost. If it is worse than the rule it would replace, stop. The record is still useful.
First show the number beside the human's decision. Then, after two clean weeks and a written sign off, let code act above a line you chose from your own data. Keep one switch that turns it off. Keep the record. You can replay it against any other model later.
Keys come from the vendor console. The API is hosted only. No weights are published and there is no self hosting.
Measured by the vendor on workflows the vendor wrote, against reference answers that were the average of two large language models. The vendor calls these "the higher end of real world gains." Independent tests show 5 to 20x on speed and cost, with slightly lower accuracy.
The whole value is that 0.9 means right about 90 percent of the time. TypeSafe trains for that. Nobody outside has published the curve. Build your own from your shadow record.
The vendor says the price may be subsidized and expects it to fall, and cannot prove otherwise. Design so you can leave: keep the record, keep the interface, pin the version.
On an invoice processing test, Jev scored 61.8 percent against 79.1 for a frontier language model. Faster and cheaper is real. Smarter is not the claim.
The vendor says it does not train on customer requests. Zero retention is an enterprise feature. Send the fields the question needs, not the whole record.
What happens when it is wrong, and who sees it? If the answer is "it is rarely wrong," walk away. If the answer is "a person sees every call it is not sure about, and here is the record," you are talking to someone who has read this page.