AI Writings · Decision Model

JEV100,000 articles. One dollar.

TypeSafe built an AI that writes nothing — it only answers which one · true or false · how many points. In exchange: every answer lands in under half a second, costs a few hundred times less than ChatGPT, never strays outside the list you gave it, and always says how sure it is.

Answers in ~0.5s Per decision ~$0.000025 Use cases 14 Read ~11 min
01 · What happened
An AI deliberately built not to write.
Your company pays ChatGPT or Claude to answer things like: which team owns this message, is this customer asking for a refund, is this application worth reading. The answer is one or two words — but you pay for a full prose machine and wait 2–10 seconds for it.
On 15 September 2026, a company called TypeSafe AI launched Jev, an AI that gave up writing entirely. It answers exactly three kinds of question: which one, true or false, how many points. In return: half a second instead of ten, and a few hundred times cheaper.
For an executive the interesting part isn't “a new model”. It's that the work you always knew you should do — score every lead, read every piece of feedback, check every line of a statement — is now worth doing.
Explainer · what a System One model is

Psychologist Daniel Kahneman split thinking in two. Type 1 is reflex: one glance at a message tells you whether it's urgent. Type 2 is deliberation: writing a proposal, weighing options.

ChatGPT and Claude are type 2 — good at writing and reasoning, but slow and expensive. Jev is type 1: fast judgement only, not a single sentence of prose. TypeSafe calls this a new category, the System One model.

A note Every figure here comes from TypeSafe's own materials; no third party has verified them. The product is in early access — open to a limited set of customers. Cost figures are estimates at list price.
02 · What it can answer
Three kinds of question. That's all.
You hand it a situation — a message, an application, an order — and ask. One question or ten at once costs almost the same.
1
Noul

True or false

“Is this customer asking for a refund?” — returns how likely it is, say 0.88.

2
Choice

Pick one

“Which team owns this?” — picks from the list you provide, up to 255 options.

3
Score

Rate it

“How frustrated is this customer?” — rates on your own scale: calm · annoyed · very angry.

Every answer comes with one number that matters more than the answer itself: confidence — how sure it is. The next section explains why that's the valuable part.
For the technical team — what one Jev call looks like

A request carries state (the raw data) and questions. Questions are scored independently and in parallel, so adding more barely costs time. Billing is on input tokens; output tokens are free.

from typesafe_sdk import TypeSafeClient, Choice, Score, Noul

client = TypeSafeClient()

res = client.system_one(
    state={"ticket": {
        "channel": "Zalo OA",
        "body": "Charged twice for last night's booking, nobody picks up the hotline.",
    }},
    questions={
        "team": Choice(
            instructions="Which team should handle this",
            criteria={
                "billing":   "Wrong amounts, refunds, invoices",
                "operations":"Food quality, service at the branch",
                "technical": "App bugs, booking failures",
            },
        ),
        "frustration": Score(
            instructions="How frustrated the customer appears",
            criteria=["Calm", "Annoyed", "Very angry"],
        ),
        "refund_requested": Noul(
            instructions="The customer is asking for a refund",
        ),
    },
)

res.answers["team"].choice              # "billing"
res.answers["team"].confidence          # 0.93
res.answers["frustration"].score        # 1.74
res.answers["refund_requested"].noul    # 0.88
03 · Strengths
Four things worth an executive's attention.

+Fast enough that nobody waits

Half a second, against 10–38 seconds for the writing models. The decision can happen while the customer is still clicking, instead of in a background job handled later.

+Cheap enough to run on everything

Roughly $0.000025 per decision. Work you previously ran on a 5% sample now runs on 100% — and that is where the real difference appears.

+It tells you how sure it is

Every answer carries a confidence number. You set the rule: below 0.5 a person looks at it, above 0.85 the system acts. Risk becomes a dial instead of a feeling.

+It never answers outside your list

It can only pick from the options you give, so it cannot invent a department that doesn't exist. Your engineers drop a whole layer of code written to catch nonsense answers.

Explainer · why confidence beats “it doesn't make things up”

A writing model can also say “I'm fairly confident”, but that's just words chosen to sound good. Jev was trained specifically so the number matches reality: when it says 0.9, it is right about 90% of the time.

Once the number is trustworthy you can write policy on top of it. That is the line between “AI assists” and “AI runs it, under control”.

04 · Weaknesses
Five places it falls short — mostly on purpose.

−It cannot tell you why

You get a number and no reasoning at all. In banking, insurance or hiring — where people can ask “why was I rejected” — that is a real wall, not an inconvenience.

−It cannot do arithmetic

Counting, adding, comparing dates — none of it is reliable. Numbers belong in your software; Jev only handles the judgement.

−Extra data makes it worse

The opposite of the usual habit of dumping everything in and letting the AI sort it out. Here you filter first, or accuracy drops.

−Customers can steer it

Someone outside your company can phrase a message so it lands in the wrong bucket. Never let a single question decide something involving money.

−Only mid-tier accuracy

In TypeSafe's own testing Jev is right 67.8% of the time, while Claude Opus 5 and GPT-5.6 Sol land around 73–74%. Six accuracy points traded for a hundred times the speed — a good deal for sorting at volume, a bad one for a single large decision.

Risks to weigh before signing anything The price is low enough that TypeSafe itself admits it cannot prove it isn't subsidised. Nobody outside the company has verified the numbers. And every time they upgrade the model, the thresholds you set have to be re-measured. If you adopt it, keep a fallback ready.
05 · Comparison
It doesn't replace ChatGPT. It replaces the place you use ChatGPT wrongly.
Criterion ChatGPT / Claude Jev Hard-coded rules
What you get back Text — it can write anything One choice or score, plus how sure it is A fixed result from the rules you wrote
How fast 2–38 seconds Under half a second Instant
Cost per 1,000 answers ~$30 ~$0.03 Near zero
Understands human language Yes, very well Yes, well enough to sort No
Can it justify itself Yes, though the reason may be invented No Yes — the rule is the reason
Best for Work that needs writing or reasoning Repeated judgement, high volume Clear rules, no context needed
Put simply: Jev takes language understanding from the hard-coded rules, and takes 99% of the invoice from ChatGPT.
06 · Business use cases
Four places where the arithmetic flips.
Figures below are estimates at list price, meant to show orders of magnitude — not a quote.
F&B · 40 branches

Sorting 8,000 messages a day to the right team

Messages from chat apps, social pages, marketplaces and the hotline land in one inbox. Each one gets five questions at once: which team, how upset the customer is, refund demanded, competitor mentioned, urgent or not. In under half a second it's in the right queue.

What actually changes: this runs on every message, so angry cases jump the queue in the first minute instead of surfacing three hours later.

$0.000025
per message
~$6
per month
Online retail

Refund approvals in three lanes

You write the policy, Jev supplies the number: low confidence → a person handles it. Medium → an AI drafts the reply and a person approves. High and order under $20 → approved automatically.

One thing is mandatory: store the situation and the numbers for every case. Jev never explains itself, so that record is all you will have later.

3 lanes
human · AI · automatic
B2B · Sales

Scoring 12,000 leads a month

Instead of asking “is this lead good” — far too vague — split it into four separate questions: right industry, budget size, does the contact have authority, how urgent. You control the weights.

Want manufacturing prioritised this quarter? Change one weight, instead of rewriting instructions for an AI and hoping.

~$0.40
to score them all / month
Any company using AI to write

A gate before anything reaches a customer

Every email, post or reply an AI writes passes through Jev first: does it leak internal information, does it promise something outside policy, is the tone right for the brand.

This is the best bargain of all: a model hundreds of times cheaper watching an expensive one, at a cost of about a tenth of a second.

~0.1s
added per send
Fourteen questions you could ask next week.
Every row is work already happening in your company: the question on the left, what the system does the moment the answer arrives on the right.
Team The question you ask Jev What happens next
Customer supportWhich team owns this message? How upset is the customer?Into the right queue, angry cases jump to the top
Customer supportIs the customer threatening to leave or to post publicly?A manager is alerted within a minute
SalesIs this lead in our target industry? Does the contact have authority?Scored, routed to the right rep, the rest into nurture
SalesAfter the call: interested, considering, or a no?CRM status updated automatically, follow-up scheduled
MarketingIs this comment praise, a complaint, a price question or spam?Auto-reply, hide spam, push price questions to sales
FinanceWhich expense account does this line belong to?Coded automatically; unsure lines go to an accountant
FinanceIs this document an invoice, a receipt or a contract?Filed in the right place, a task is opened
HRDoes this application meet the three required criteria?First-pass filter; humans only read what survives
HRDoes this internal report show signs of something serious?Escalated straight to senior management
OperationsWhat kind of incident is this, and does it block production?Ticket opened for the right team with a priority
Retail · F&BIs this 1-star review about the food, the service or the delivery?Sent to the right branch and department, with weekly stats
LogisticsIs this complaint about a late delivery, damaged packaging or a wrong item?The correct compensation process starts without asking again
Real estate · ServicesIs this enquiry a real buyer or another agent pitching?Filtered before it reaches a salesperson
Legal · RiskDoes this outgoing message promise anything outside policy?Held back for a human to approve before it goes out
The same approach covers sorting 30,000 bank-statement lines a month into the right expense accounts, screening CVs criterion by criterion, or tagging survey responses. The common thread: repetitive work, answers that fit a known list, and a job previously dropped because it wasn't worth the money.
07 · Can a normal company use it
You still need someone technical — but less than you'd think.
Plainly: Jev has no app to click, no drag-and-drop screen. It is an API — an address software calls. Someone has to wire it up.
But this is not a months-long AI project. One question is a few lines of code. The hard part isn't the technical part.

01Through whoever does your automation

The common workflow tools — n8n, Make, Zapier, even Google Apps Script — can all call an API. Whoever builds your automations can do this. The first use case usually takes days, not months.

02Through infrastructure you already pay for

Jev is available on Cloudflare. If your site or app already runs there, this is switching on one more service rather than onboarding a new vendor.

03Through your software team or partner

If you already run a CRM, a sales system or a call centre, it goes in at exactly one point: the moment data arrives, before a person touches it.

Four things the business side has to prepare

  • Pick exactly one task — the most repetitive one, where you can count the daily volume. Not five at once
  • Write down the answer list with a description for each option — this is simply how your company already sorts the work
  • Pull 200–500 past cases whose outcome you already know, to measure how often the model agrees
  • Set the threshold and run in parallel for two weeks: the machine decides, people also decide, compare weekly before going live
In short: the tool sits on the technical side, but the problem sits on the business side. The hard part isn't writing code — it's stating the question and the answer list clearly. That is an operator's job, not a programmer's.
08 · Adding it to an AI system you already run
This is where Jev moves the numbers a business actually watches.
Jev is a developer's tool, yes. But what it changes is how your AI system is arranged — and that arrangement decides four numbers the board looks at: monthly cost, response time, how much runs without a person, and how much risk you can afford to take.
Tier 1 · Software

Rules and numbers

Everything with a definite right answer: billing, permissions, running the loop.

→
Tier 2 · Jev

Fast judgement

Sorting, routing, scoring — where hard rules can't capture the nuance.

→
Tier 3 · Writing AI

When words are needed

Drafting, summarising, analysing odd cases. Only on the small remainder.

→
Tier 4 · People

Hard and risky cases

The final call — and the source of truth for re-measuring quality each month.

Most AI systems today do the opposite: one expensive model owns everything — reading, judging, writing, deciding — and the engineering team patches whatever it breaks. Rearranged into four tiers, the expensive model steps back to what it is genuinely best at: writing.
Six places to insert it into a system already running an LLM.
No rebuild required. Each of these is a separate insertion point you can do one at a time.

01A gate at the front door

Jev sorts first; the LLM only runs on what genuinely needs writing. This is where most of the money is: today the majority of LLM calls exist to retrieve a one-word judgement.

02Pick the model by difficulty

Jev rates whether a case is easy or hard, low or high risk — and the system routes it to the cheap model or the expensive one. Quality holds where it matters, cost drops everywhere else.

03Filter the data before the LLM sees it

Only the relevant part goes into the prompt. A shorter prompt is both cheaper and more accurate — the more irrelevant material an LLM gets, the more it drifts.

04Guard the output

Everything the LLM writes gets checked before it ships. Cheap enough to run on 100% of content, instead of the few per cent spot-check most teams do today.

05Score every case, not a sample

Checking whether AI answers were any good used to require an expensive LLM as the judge, so teams graded 2–5% of cases. Now grade all of them — quality becomes a weekly number, and a drop shows up immediately.

06Cut the agent's loop

Agents ask an LLM at every fork: keep going or stop, which tool to call, is this done. Move those forks to Jev and the agent runs many times faster, drifts less, and stops burning money on redundant thinking.

Before and after, on a support system handling 8,000 cases a day.
Illustrative estimates at list price, assuming 20% of cases genuinely need an AI-drafted reply.
Metric LLM only Jev in front, LLM behind
LLM calls per day8,000 — every case1,600 — only cases needing words
Model cost per day~$240~$48
Model cost per month~$7,200~$1,450
Time to sort one case4–10 secondsUnder 0.5 seconds
Malformed output to retryYes — needs code to handle itEffectively none at the sorting layer
Share of cases quality-checkedA 2–5% sample100%
Cases running without a personFew — no trustworthy number to threshold onEverything above your confidence bar
Where those six changes reach the business.

What the rearrangement makes possible

  • AI cost per order, per ticket, per customer drops by an order of magnitude. Products that never worked at the price — cheap tiers, small customers, price-sensitive markets — can now clear the bar. That's a business-model change, not a technology one.
  • AI moves from a background job to something the customer sees happen. Competitors on plain LLMs need 10–30 seconds, so they must answer later; you answer while the customer is still on the screen. Same feature, entirely different experience.
  • Growth stops requiring headcount in a straight line. Easy cases clear themselves above the threshold; people handle the hard ones. Doubling orders no longer means doubling the operations team.
  • You can put AI in front of customers. With a guard running on 100% of content at negligible cost, brand risk is controlled — instead of keeping AI internal because it felt safer.
  • Less vendor lock-in. Once judgement lives outside any one model's prompt, switching LLM later means replacing a tier rather than rewriting the system. Your negotiating position changes with it.
  • Quality becomes a number instead of a feeling. Grading 100% weekly gives leadership something to decide on: where to expand, where to tighten, where people still have to stay in the loop.
When it isn't worth it yet This is a question of scale. If your system makes a few hundred AI calls a day, the savings won't cover the cost of rearranging and re-measuring. The threshold worth considering usually starts around a few thousand calls a day — or the moment response time becomes something customers notice.
09 · What's still missing
Four things that decide whether it lasts.

01A trail you can show

It doesn't need to write reasons — just point at which part of the situation tipped the answer. Enough to build a record when a customer or a regulator asks.

02Numbers and dates

If a later version handles them, a whole slice of finance work — reconciliation, spotting anomalies — opens up at once.

03Learning your own categories

Today you can't feed real outcomes back so it converges on how your company sorts things. That's the line between a handy tool and lasting infrastructure.

04Commitments on versions and price

Every version change forces you to re-measure your thresholds. That needs a clear roadmap, plus numbers verified by someone other than the vendor.

10 · Verdict
When to use it, and when never to.

Use it for

  • Work repeated thousands of times a day
  • Answers that fit a list you can write down in advance
  • Decisions that must land in under a second, while someone waits
  • Work previously skipped because it wasn't worth the money
  • Checking AI-written content before it reaches a customer

Don't use it for

  • Anything that has to produce words for a person to read
  • Decisions you must justify to a customer or a regulator
  • Anything involving counting, maths or dates
  • Large decisions made once and never repeated
  • Running automatically with no threshold, no record and no review
The most valuable thing about Jev isn't the speed or the price — it's that it tells the truth about how sure it is. A system willing to say “I'm only 40% sure” is far more useful than one that always answers smoothly, because only the first gives you somewhere to put the brake.
If only one line survives: don't replace your current AI with Jev — let it write, and hand the repeated judgement to something hundreds of times cheaper.
Appendix
The terms, in plain words.
System One model
A model that only makes fast decisions and writes nothing. The opposite of a ChatGPT-style model.
LLM
The general name for writing AIs such as ChatGPT, Claude and Gemini.
State
The situation you hand the AI: a message, an application, an order.
Noul · Choice · Score
The three question types: true/false · pick one · rate it.
Confidence
How sure the AI is, from 0 to 1. What you build automation thresholds on.
Type-safe
It can only answer inside the shape you defined — never something off the list.
Hallucination
An AI inventing convincing but false information. Jev blocks the off-list kind, but can still pick the wrong option.
Early access
Open to a limited set of customers before a general release.
Benchmark
A standard test set used to compare accuracy across models.
← Back to AI Writings