Psychologist Daniel Kahneman split thinking in two. Type 1 is reflex: one glance at a message tells you whether it's urgent. Type 2 is deliberation: writing a proposal, weighing options.
ChatGPT and Claude are type 2 — good at writing and reasoning, but slow and expensive. Jev is type 1: fast judgement only, not a single sentence of prose. TypeSafe calls this a new category, the System One model.
True or false
“Is this customer asking for a refund?” — returns how likely it is, say 0.88.
Pick one
“Which team owns this?” — picks from the list you provide, up to 255 options.
Rate it
“How frustrated is this customer?” — rates on your own scale: calm · annoyed · very angry.
For the technical team — what one Jev call looks like
A request carries state (the raw data) and questions. Questions are scored independently and in parallel, so adding more barely costs time. Billing is on input tokens; output tokens are free.
from typesafe_sdk import TypeSafeClient, Choice, Score, Noul client = TypeSafeClient() res = client.system_one( state={"ticket": { "channel": "Zalo OA", "body": "Charged twice for last night's booking, nobody picks up the hotline.", }}, questions={ "team": Choice( instructions="Which team should handle this", criteria={ "billing": "Wrong amounts, refunds, invoices", "operations":"Food quality, service at the branch", "technical": "App bugs, booking failures", }, ), "frustration": Score( instructions="How frustrated the customer appears", criteria=["Calm", "Annoyed", "Very angry"], ), "refund_requested": Noul( instructions="The customer is asking for a refund", ), }, ) res.answers["team"].choice # "billing" res.answers["team"].confidence # 0.93 res.answers["frustration"].score # 1.74 res.answers["refund_requested"].noul # 0.88
+Fast enough that nobody waits
Half a second, against 10–38 seconds for the writing models. The decision can happen while the customer is still clicking, instead of in a background job handled later.
+Cheap enough to run on everything
Roughly $0.000025 per decision. Work you previously ran on a 5% sample now runs on 100% — and that is where the real difference appears.
+It tells you how sure it is
Every answer carries a confidence number. You set the rule: below 0.5 a person looks at it, above 0.85 the system acts. Risk becomes a dial instead of a feeling.
+It never answers outside your list
It can only pick from the options you give, so it cannot invent a department that doesn't exist. Your engineers drop a whole layer of code written to catch nonsense answers.
A writing model can also say “I'm fairly confident”, but that's just words chosen to sound good. Jev was trained specifically so the number matches reality: when it says 0.9, it is right about 90% of the time.
Once the number is trustworthy you can write policy on top of it. That is the line between “AI assists” and “AI runs it, under control”.
−It cannot tell you why
You get a number and no reasoning at all. In banking, insurance or hiring — where people can ask “why was I rejected” — that is a real wall, not an inconvenience.
−It cannot do arithmetic
Counting, adding, comparing dates — none of it is reliable. Numbers belong in your software; Jev only handles the judgement.
−Extra data makes it worse
The opposite of the usual habit of dumping everything in and letting the AI sort it out. Here you filter first, or accuracy drops.
−Customers can steer it
Someone outside your company can phrase a message so it lands in the wrong bucket. Never let a single question decide something involving money.
−Only mid-tier accuracy
In TypeSafe's own testing Jev is right 67.8% of the time, while Claude Opus 5 and GPT-5.6 Sol land around 73–74%. Six accuracy points traded for a hundred times the speed — a good deal for sorting at volume, a bad one for a single large decision.
| Criterion | ChatGPT / Claude | Jev | Hard-coded rules |
|---|---|---|---|
| What you get back | Text — it can write anything | One choice or score, plus how sure it is | A fixed result from the rules you wrote |
| How fast | 2–38 seconds | Under half a second | Instant |
| Cost per 1,000 answers | ~$30 | ~$0.03 | Near zero |
| Understands human language | Yes, very well | Yes, well enough to sort | No |
| Can it justify itself | Yes, though the reason may be invented | No | Yes — the rule is the reason |
| Best for | Work that needs writing or reasoning | Repeated judgement, high volume | Clear rules, no context needed |
Sorting 8,000 messages a day to the right team
Messages from chat apps, social pages, marketplaces and the hotline land in one inbox. Each one gets five questions at once: which team, how upset the customer is, refund demanded, competitor mentioned, urgent or not. In under half a second it's in the right queue.
What actually changes: this runs on every message, so angry cases jump the queue in the first minute instead of surfacing three hours later.
Refund approvals in three lanes
You write the policy, Jev supplies the number: low confidence → a person handles it. Medium → an AI drafts the reply and a person approves. High and order under $20 → approved automatically.
One thing is mandatory: store the situation and the numbers for every case. Jev never explains itself, so that record is all you will have later.
Scoring 12,000 leads a month
Instead of asking “is this lead good” — far too vague — split it into four separate questions: right industry, budget size, does the contact have authority, how urgent. You control the weights.
Want manufacturing prioritised this quarter? Change one weight, instead of rewriting instructions for an AI and hoping.
A gate before anything reaches a customer
Every email, post or reply an AI writes passes through Jev first: does it leak internal information, does it promise something outside policy, is the tone right for the brand.
This is the best bargain of all: a model hundreds of times cheaper watching an expensive one, at a cost of about a tenth of a second.
| Team | The question you ask Jev | What happens next |
|---|---|---|
| Customer support | Which team owns this message? How upset is the customer? | Into the right queue, angry cases jump to the top |
| Customer support | Is the customer threatening to leave or to post publicly? | A manager is alerted within a minute |
| Sales | Is this lead in our target industry? Does the contact have authority? | Scored, routed to the right rep, the rest into nurture |
| Sales | After the call: interested, considering, or a no? | CRM status updated automatically, follow-up scheduled |
| Marketing | Is this comment praise, a complaint, a price question or spam? | Auto-reply, hide spam, push price questions to sales |
| Finance | Which expense account does this line belong to? | Coded automatically; unsure lines go to an accountant |
| Finance | Is this document an invoice, a receipt or a contract? | Filed in the right place, a task is opened |
| HR | Does this application meet the three required criteria? | First-pass filter; humans only read what survives |
| HR | Does this internal report show signs of something serious? | Escalated straight to senior management |
| Operations | What kind of incident is this, and does it block production? | Ticket opened for the right team with a priority |
| Retail · F&B | Is this 1-star review about the food, the service or the delivery? | Sent to the right branch and department, with weekly stats |
| Logistics | Is this complaint about a late delivery, damaged packaging or a wrong item? | The correct compensation process starts without asking again |
| Real estate · Services | Is this enquiry a real buyer or another agent pitching? | Filtered before it reaches a salesperson |
| Legal · Risk | Does this outgoing message promise anything outside policy? | Held back for a human to approve before it goes out |
01Through whoever does your automation
The common workflow tools — n8n, Make, Zapier, even Google Apps Script — can all call an API. Whoever builds your automations can do this. The first use case usually takes days, not months.
02Through infrastructure you already pay for
Jev is available on Cloudflare. If your site or app already runs there, this is switching on one more service rather than onboarding a new vendor.
03Through your software team or partner
If you already run a CRM, a sales system or a call centre, it goes in at exactly one point: the moment data arrives, before a person touches it.
Four things the business side has to prepare
- Pick exactly one task — the most repetitive one, where you can count the daily volume. Not five at once
- Write down the answer list with a description for each option — this is simply how your company already sorts the work
- Pull 200–500 past cases whose outcome you already know, to measure how often the model agrees
- Set the threshold and run in parallel for two weeks: the machine decides, people also decide, compare weekly before going live
Rules and numbers
Everything with a definite right answer: billing, permissions, running the loop.
Fast judgement
Sorting, routing, scoring — where hard rules can't capture the nuance.
When words are needed
Drafting, summarising, analysing odd cases. Only on the small remainder.
Hard and risky cases
The final call — and the source of truth for re-measuring quality each month.
01A gate at the front door
Jev sorts first; the LLM only runs on what genuinely needs writing. This is where most of the money is: today the majority of LLM calls exist to retrieve a one-word judgement.
02Pick the model by difficulty
Jev rates whether a case is easy or hard, low or high risk — and the system routes it to the cheap model or the expensive one. Quality holds where it matters, cost drops everywhere else.
03Filter the data before the LLM sees it
Only the relevant part goes into the prompt. A shorter prompt is both cheaper and more accurate — the more irrelevant material an LLM gets, the more it drifts.
04Guard the output
Everything the LLM writes gets checked before it ships. Cheap enough to run on 100% of content, instead of the few per cent spot-check most teams do today.
05Score every case, not a sample
Checking whether AI answers were any good used to require an expensive LLM as the judge, so teams graded 2–5% of cases. Now grade all of them — quality becomes a weekly number, and a drop shows up immediately.
06Cut the agent's loop
Agents ask an LLM at every fork: keep going or stop, which tool to call, is this done. Move those forks to Jev and the agent runs many times faster, drifts less, and stops burning money on redundant thinking.
| Metric | LLM only | Jev in front, LLM behind |
|---|---|---|
| LLM calls per day | 8,000 — every case | 1,600 — only cases needing words |
| Model cost per day | ~$240 | ~$48 |
| Model cost per month | ~$7,200 | ~$1,450 |
| Time to sort one case | 4–10 seconds | Under 0.5 seconds |
| Malformed output to retry | Yes — needs code to handle it | Effectively none at the sorting layer |
| Share of cases quality-checked | A 2–5% sample | 100% |
| Cases running without a person | Few — no trustworthy number to threshold on | Everything above your confidence bar |
What the rearrangement makes possible
- AI cost per order, per ticket, per customer drops by an order of magnitude. Products that never worked at the price — cheap tiers, small customers, price-sensitive markets — can now clear the bar. That's a business-model change, not a technology one.
- AI moves from a background job to something the customer sees happen. Competitors on plain LLMs need 10–30 seconds, so they must answer later; you answer while the customer is still on the screen. Same feature, entirely different experience.
- Growth stops requiring headcount in a straight line. Easy cases clear themselves above the threshold; people handle the hard ones. Doubling orders no longer means doubling the operations team.
- You can put AI in front of customers. With a guard running on 100% of content at negligible cost, brand risk is controlled — instead of keeping AI internal because it felt safer.
- Less vendor lock-in. Once judgement lives outside any one model's prompt, switching LLM later means replacing a tier rather than rewriting the system. Your negotiating position changes with it.
- Quality becomes a number instead of a feeling. Grading 100% weekly gives leadership something to decide on: where to expand, where to tighten, where people still have to stay in the loop.
01A trail you can show
It doesn't need to write reasons — just point at which part of the situation tipped the answer. Enough to build a record when a customer or a regulator asks.
02Numbers and dates
If a later version handles them, a whole slice of finance work — reconciliation, spotting anomalies — opens up at once.
03Learning your own categories
Today you can't feed real outcomes back so it converges on how your company sorts things. That's the line between a handy tool and lasting infrastructure.
04Commitments on versions and price
Every version change forces you to re-measure your thresholds. That needs a clear roadmap, plus numbers verified by someone other than the vendor.
Use it for
- Work repeated thousands of times a day
- Answers that fit a list you can write down in advance
- Decisions that must land in under a second, while someone waits
- Work previously skipped because it wasn't worth the money
- Checking AI-written content before it reaches a customer
Don't use it for
- Anything that has to produce words for a person to read
- Decisions you must justify to a customer or a regulator
- Anything involving counting, maths or dates
- Large decisions made once and never repeated
- Running automatically with no threshold, no record and no review