Pyyan / Boxes / The model that answers yes, and nothing else

Explainer

The model that answers yes, and nothing else

Jev and Laya do not write. They take a state and a typed question and hand back an answer with a probability, in about a tenth of a second. What this class of model is for, what it costs, which of the claims about it hold up, and when a language model is still the right tool.

Pyyan24 September 202612 min read

Ask a language model whether a refund request is fraudulent and it will write you a paragraph. You did not want a paragraph. You wanted a true or a false, a number saying how sure it was, and an answer fast enough to sit inside a request handler.

That gap is the whole product. In one week in September 2026, two models shipped that do only the second thing: Jev, from a San Francisco company called TypeSafe AI, and Laya, from a four-person company in Kasaragod, Kerala. Neither writes. Both take a state and a typed question and return a typed answer with a probability, in roughly a tenth of a second.

#What a decision model is

Three terms first, because the rest of this page leans on them.

A state is whatever you already have: an email, a support ticket, a JSON object, a chunk of a document. You do not prompt it into a sentence. You pass it as it is.

A typed question is a question with a declared answer shape. Both models use the same three:

Type You ask You get back
noul yes or no one probability
choice pick one option the pick, and odds for each
score place on a scale a score and a confidence

Non-autoregressive means the model does not write its answer one token at a time. A language model generates left to right, and each token waits for the one before it, which is where the latency comes from. A decision model scores every candidate answer in parallel in a single forward pass. There is no sentence to produce, so there is no queue.

That is the whole architecture story. What you get out is a value, not prose:

state      "Package never arrived. Third
            claim this month. Same address."

questions  is_refund_fraud   noul
           route             choice
           severity          score 1-5

answers    is_refund_fraud   true    0.97
           route             fraud   0.81
           severity          4 of 5  0.64

Your code branches on those numbers. The useful pattern is not "let the model decide" but let the model decide when it is sure, and send the rest to a person: if the probability is above 0.95, act; otherwise escalate. A threshold you used to hard-code becomes a judgement with a confidence attached to it.

#Jev, and the two things everyone gets wrong about it

Jev was released on 15 September 2026 in limited early access by TypeSafe AI, founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. Almeida spent four years at OpenAI, on RLHF, InstructGPT, ChatGPT and GPT-4. The company raised a $40m seed led by DCVC at a $200m valuation, and calls Jev the first "System One" model, after the fast, intuitive thinking in Daniel Kahneman's split.

The published numbers, all of them TypeSafe's own:

Jev
Latency 70 to 500 ms, most near 100
Price $0.042 per M input, output free
Rate limit 250k tokens/sec, 1,200 req/min
Claimed 40 to 200x faster, 40 to 400x cheaper
Peak claim 193.6x faster, 444.6x cheaper
Trained on synthetic data only
Published no paper, no weights, no size

The training method has a name, Reinforcement Learning for Calibrated Decisions, and no paper behind it.

Two claims travel with Jev that do not survive checking.

It is not rule-based. It is a transformer, trained on synthetic data, that returns probabilities. The confusion is understandable, because the thing it replaces is a rule: the if statement you wrote that was too brittle, or the classifier you never got around to training. TypeSafe's own framing is an AI-native if. A rule engine executes conditions you wrote. Jev estimates an answer you did not write, and tells you how confident it is. Those are opposites in every way that matters for testing, auditing and debugging.

The free API on Vercel is neither unlimited nor open-ended. Vercel's AI Gateway listed Jev on 16 September with the line "free on AI Gateway until September 25th". That is nine days, ending tomorrow at the time of writing, 24 September 2026. It needs a card on file, a team without a payment method gets customer_verification_required, and the free tier is rate-limited well below the paid one, tightly enough that a handful of questions in a row get throttled. Free for a week, with a card, at a low ceiling, is a good way to try something. It is not an unlimited API.

#Laya, and where it came from

Four days after Jev, on 19 September 2026, Convai Innovations published Laya under Apache 2.0, with the weights on Hugging Face and no API to subscribe to. Convai is based in Kasaragod, Kerala, founded in 2021 by Nandakishor M with Archana M and Anjali M, and is on the Kerala Startup Mission register. Its earlier product is AI for Cardio, which reads ECG traces in real time. Nandakishor was previously a principal investigator at the IIT Palakkad Technology IHUB Foundation.

Laya is not a repackaged language model. It is an encoder with decision heads bolted on, which is the older and much cheaper end of the same family tree:

Checkpoint Backbone Size Context
laya ModernBERT-large 421M 512
-multilingual mmBERT-base 322M 1,024+
-typed-decisions ModernBERT-large 421M 1,024

The first is English, for guardrails and email triage. The second covers 100+ languages and runs about 2.2 times faster. The third is tuned for four typed-decision workflows.

Measured on one T4, a GPU you can rent for pocket money: 32.8 ms for a single multilingual question, 39.5 ms in English, and 7.2 ms per question when you batch them. The comparison Convai publishes, against Jev 1.13.0:

Laya Jev 1.13.0
typed-decisions 0.766 0.727
AG News 0.950 0.910
DAIR Emotion 0.595 0.480
Latency 32.8 ms 236 to 276 ms
Calibration error 0.081 0.246

Lower is better on the last row only.

Every one of those numbers is Convai's, measured by Convai. That is not an accusation, it is the state of the evidence: TypeSafe's numbers are TypeSafe's too, and nobody independent has run either model against the other. Treat the table as a claim with a method attached, which is more than most launches offer, and less than a result.

The lineage is worth knowing, because it explains why a four-person company shipped this in a week rather than a year. Nandakishor's own arXiv papers run straight into it: a reinforcement learning model for predicting sales conversion in real time (March 2025), and confidence-aware routing that decides before generating whether to answer locally, retrieve, escalate to a bigger model, or hand to a human (September 2025). Laya is the third step on a path he was already walking.

#The two of them, side by side

Jev Laya
Who TypeSafe AI, SF Convai, Kasaragod
Released 15 Sep 2026 19 Sep 2026
Weights none Apache 2.0
Runs hosted API your machine
Size not published 322M to 421M
Price $0.042 per M in electricity
Languages not published 100+
Evidence vendor numbers numbers, paper, weights

The honest summary is that they are the same idea at two prices. If the decision is small, frequent and private, a 322M model on your own box answering in 33 ms is hard to argue with. If you do not want to run anything, Jev is a URL.

A third exists: meraGPT's Decider 1, which speaks the same SDK if you change the base URL. The interface is becoming a category faster than the models are.

#How you actually call it

Three examples, in the shape each vendor publishes. The pattern is identical every time: hand over a state, hand over named questions with their types, read typed answers back.

#Laya, on your own machine

No key, no account, no network. The first call downloads the checkpoint and picks which of the three to use.

pip install laya
from laya import Router

router = Router()

state = """Billed twice for March. Refund the duplicate
today or we cancel the plan."""

questions = {
    "department": {
        "type": "choice",
        "instructions": "Who should handle this?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, errors",
            "other": "everything else",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["not urgent", "soon", "blocking"],
    },
    "churn_risk": {
        "type": "noul",
        "instructions": "Do they threaten to leave?",
    },
}

r = router.predict(state, questions)

r["answers"]["department"]["choice"]   # "billing"
r["answers"]["churn_risk"]["noul"]     # 0.892
r["routing"]["model"]                  # "english"

Note criteria. For a choice it is a dictionary, because each option needs saying what it means. For a score it is an ordered list, lowest first. For a noul you can leave it out. Those rubrics are the whole prompt engineering surface of this class of model, and they are where accuracy comes from.

#Jev, from Python

Same idea, an API key instead of a download, and one extra move: a noul can carry a definition of what true and false each mean, which is the neatest way to pin down a vague question.

from typesafe_sdk import (
    TypeSafeClient, Noul, NoulCriteria, Score,
)

jev = TypeSafeClient()   # reads TYPESAFE_API_KEY

r = jev.system_one(
    state={"page": page_text, "query": user_query},
    questions={
        "answers_query": Noul(
            instructions="Does the page answer it?",
            criteria=NoulCriteria(
                true="Explains what was asked",
                false="Related, but no answer",
            ),
        ),
        "frustration": Score(
            instructions="How annoyed is the writer?",
            criteria=["calm", "annoyed", "angry"],
        ),
    },
)

r.answers["answers_query"].noul       # 0.97
r.answers["frustration"].score        # 2.1
r.answers["frustration"].confidence   # 0.64
r.usage.output_tokens                 # 0, always

That last line is not a curiosity. Output tokens are always zero because nothing is generated, which is why output is free and why the latency has no tail: there is no sentence whose length you are waiting on.

#The pattern worth copying

Neither example above is useful on its own. This is:

ANSWER, ESCALATE = 0.95, 0.70

p = r["answers"]["churn_risk"]["noul"]

if p >= ANSWER:
    page_retention_team(ticket)
elif p >= ESCALATE:
    queue_for_human(ticket, note=f"maybe churn {p:.2f}")
else:
    route_normally(ticket)

Two thresholds, not one. The top band acts, the middle band asks a person, the bottom band ignores it. You tune the two numbers against labelled examples once and then the rubric, not the code, is what you edit. This is the entire promise of the category: the confidence score is the product, and a model that is well calibrated lets you spend human attention only in the middle band.

#Jev through the Vercel gateway, in TypeScript

If you are already on the AI SDK, Jev arrives as a verb rather than a chat call. Worth knowing: the type called noul in Python is called boolean here, and the confidence comes back beside the answers rather than inside them.

import { experimental_evaluate as evaluate } from 'ai';

const { answers, providerMetadata } = await evaluate({
  model: 'typesafe-ai/jev',
  state: { prompt, repoSummary },
  questions: {
    difficulty: {
      type: 'score',
      instructions: 'How much reasoning is needed?',
      criteria: ['simple', 'complex', 'very hard'],
    },
    needsWeb: {
      type: 'boolean',
      instructions: 'Does this need external docs?',
    },
  },
});

const conf =
  providerMetadata?.typesafe?.confidence?.difficulty;

A router in front of a coding agent is the obvious use: cheap model for simple, frontier model for very hard, web search only when the second answer says so. The decision costs about a tenth of a cent per thousand calls and saves a frontier call every time it says simple.

#What neither of them can do

This is the part the launch posts skip, and it is short and sharp:

  • Arithmetic and counting. Reliably wrong. Do the sum in your code.
  • Dates, measurements, indirect references. "The Tuesday after the one we discussed" is not a question this class of model answers.
  • Negation. Questions phrased in the negative confuse it. Ask the positive one and flip the result yourself.
  • Anything generative. No writing, no code, no images, no audio. If you need a candidate produced rather than chosen, you need a different model.
  • Long, noisy state. Irrelevant detail in the state degrades accuracy, a failure people have taken to calling context rot. Pass the part that matters.
  • Being right. A constrained output cannot hallucinate a citation, because it cannot write one. It can still be confidently wrong, and the probability it returns is only as good as its calibration. Schema safety is not correctness.

#Is this new?

The loudest objection to Jev, from r/LocalLLaMA on the day, was that it is not: "Jev isn't new tech. Its marketing targets people who think AI started with LLMs." Another thread said the quiet part with more nuance: it is a good generalist for when you do not know what you need, and for almost any specific use case there is a smaller, faster local alternative.

Both are fair. Text classifiers are as old as the field, BERT-style encoders have been doing this since 2018, and Laya's own backbones are exactly that. What is actually new is narrower and still worth something:

  1. Zero-shot typed questions. You do not train a classifier per task and you do not collect a labelled set first. You write the question and the rubric.
  2. Calibration as a product feature. The number that decides whether a human sees the case is the thing being sold, not a side effect.
  3. One interface over arbitrary state. The same call answers a boolean, a routing choice and a severity score against the same blob of text.

The reason it arrived now is cost. Two years of building agents taught everyone that the expensive part is not the clever answer at the end, it is the thousand small judgements on the way there, each one currently costing a full language model call. Hacker News put it plainly on 21 September, with 318 points behind it: OpenAI is well positioned to fast-follow Jev. Of course it is. The interface is the invention, and interfaces get copied.

#How to decide which, or neither

  • The decision is bounded and you can write the rubric: a decision model.
  • You need prose, code, or a plan: a language model. This is not a substitute.
  • You need an audit trail of reasoning: neither. You get a number, not a because.
  • It must run on your hardware, on your data: Laya today, since Jev has no weights.
  • You want nothing to operate: Jev, priced per million tokens.
  • You are doing arithmetic, dates or counting: ordinary code, as always.

#What to watch

Three things would change this page.

An independent benchmark. Right now both vendors grade their own homework, and JevBench is run by a third party nobody has audited either. The first neutral evaluation of calibration across both models is the number that matters.

The price of the frontier. On 22 September Anthropic and OpenAI both cut flagship prices, GPT-6 Luna landing at $0.10 per million input tokens. Decision models justify themselves on a cost gap; that gap narrowed the week they launched.

Whether the labs ship one. A typed-decision endpoint from OpenAI or Anthropic would take the interface and leave the startups with the part that is genuinely theirs, which is latency and self-hosting. Laya, being 322M parameters and Apache 2.0 on your own machine, is the one that survives that.

Updated 1 October 2026. They shipped one. At DevDay on 29 September, ten days after this page was written, OpenAI announced a Decisions API in limited preview: an endpoint that constrains a model to a set of predefined answer choices for fast classification, which is the interface described at the top of this page. No price, no latency figure and no benchmark were published with it, so the two things that made the category worth a page, cost and response time, are still unmeasured at OpenAI. The part of the prediction that holds is the rest of the sentence: a 322M model under Apache 2.0 on your own hardware is not something an endpoint can take away from you.

A day later, on 30 September, Inception put Mercury Decide on OpenRouter, free for early access: the same typed-question interface, served by a diffusion model, with the probability read out of the model rather than written as text, so output tokens cost nothing. That is four companies selling this interface within a fortnight of the first. The interface was the invention, as this page said, and interfaces get copied faster than models do.

Every figure on this page is dated 24 September 2026, except the update above, and comes from the companies, their model cards, their papers or the Vercel changelog, each linked at its first use. No independent benchmark of either model existed at the time of writing. If one appears, this page is wrong until it is updated.