决策模型在 2026 年 9 月 15 日 Jev 发布之后进入高速爆发期:两周内出现 100 多个开源复刻、两套公开评测、多家大厂的对标产品,以及 arXiv 上每天新增的相关论文。信息增长的速度,已经超过任何人靠追更能跟上的程度。
这里做的事情是「汇总与统一」:把散落在 Hugging Face、arXiv、厂商文档、第三方目录里的模型、评测、论文和工程经验收进同一个数据库,用同一套口径呈现,并保留每一条数据的原始出处。目标不是做一个更好的榜单,而是让任何人在十分钟内建立对这类模型的完整认知,并在需要时直达一手资料。
「决策模型」或许叫直觉模型更贴切。
观点
它并不做决策——它做的是直觉判断。更贴切的名字是「直觉模型」。
决策包含权衡、后果与责任;而这类模型做的事情是:在你给定的选项上,凭已经内化的知识给出一个概率分布,快速指一个「最像对」的答案。它不推演、不比较代价、不承担后果。就过程而言,它更接近人「拍脑袋」的那一下,而不是「想清楚之后再做决定」。
按双过程理论的说法,它对应的是 System 1——快速、直觉、省力;而不是 System 2 那种慢速、可追溯的推理。TypeSafe 自己把这个家族命名为 System One,其实比「决策模型」这个中文译法更准确。
名字重要,因为它决定了人们期待什么。叫「决策模型」,会让人以为它可以替你把事决定下来;叫「直觉模型」,你就知道它给的是一个先验判断,该不该照它执行,仍然是你的责任。
它让本来就难解释的大模型,更难解释了
观点
拿掉了生成过程,也就拿掉了唯一能被人读的那部分。
大语言模型的可解释性本来就是一笔糊涂账:注意力权重、神经元激活、思维链文本,没有一个是可靠的因果解释。但至少思维链给了人一个读得下去的东西——哪怕它可能只是事后编出来的理由。
决策模型把这一层直接删掉了:不生成任何文本,只返回一个概率分布。结果是更难解释。概率可以校准(说 0.9 就真的九成对),但「为什么是 0.9」无处可问——没有文本、没有中间态、没有可回放的过程。
所以工程上的合理位置是:低风险、可回滚、有人兜底的判断点。用它的概率做分流与阈值门控,把不确定的那部分交给人或交给更强的模型,而不是把它当作不可质疑的结论。
这个用法早就有,代价也很早就暴露了
工程经验
要求模型只输出 ABCD,确实不到一秒;但没有了 think,质量差很多。
在 Jev 出现之前,很多团队在业务里已经用过同类做法:把要求和答案结构写进上下文,要求模型只输出一个选项字母。这样做的收益非常直接——不需要解析自然语言、不需要处理格式漂移、延迟压到亚秒级、成本几乎可以忽略。
但问题同样直接:没有了思考过程,回答质量明显更差。同一道判断,让模型先写出理由再给标签,往往比直接吐标签更准;把推理过程拿掉,等于主动放弃了模型在困难样本上的那部分能力。
这正是决策模型的取舍:用质量上界换速度与成本。它适合那些本来就不需要深想的判断(这条消息该归到哪个类别、这个截图里有没有异常、这段话是否违规);而不适合把一道需要推理的问题硬压成一次前向。选型时先问一句:这件事人来做需要想吗?需要想,就别交给直觉模型。
为什么是现在爆发
趋势判断
需求一直在,缺的是行业共识;接口一旦统一,剩下的就是工程量。
这类需求并不新鲜,过去十年里以各种形式存在过:文本分类、意图识别、内容审核、rerank、规则引擎加模型打分。真正缺的不是技术,而是行业共识——可做的事情太多、收益分散在无数个具体场景里,没有人愿意为一个标准投入。
Jev 做对的一件事,是把接口形态固定下来:state + 类型化问题(Choice / Noul / Score)→ 概率分布,输出免费、按输入计费。标准一旦清晰,生态就会自己长出来:两周内 100 多个开源复刻、两套公开评测(S1MB 与 Decision Index)、Cloudflare 开源 Clef 并配套 RL 微调、OpenAI 与 Perplexity 各自推出 Decisions 端点。
共识形成之后,剩下的就是工程与数据——所以接下来会进入快车道。本站的存在意义也在于此:这段路上每天都有新东西,需要有人把它收拢、对齐、留档。
各家实现差异很大,本质是同一件事
观点
都是利用大语言模型已经内化的知识,做一次快速的、可能不那么精确的判断。
从外面看,这些模型差异极大:有在冻结的 27B 主干上加一个 head 的,有在 400M 编码器上重头训练的;有读标签 token 的 logit 的,有把选项拼进输入做成对打分的;有单次前向的,也有走受限解码的。尺寸从 17M 到 35B,延迟从 5 毫秒到 500 毫秒。
但如果问「它凭什么能判断」,答案只有一个:底模在预训练阶段已经把这些判断所需的知识内化进去了。决策模型做的事情,是把这些知识在一次前向里读出来,而不是重新学一遍。这也解释了为什么:
换句话说,架构决定的是「读得多准」,数据决定的是「读得到什么」。具体流派见下方「实现方式探索」。
- 小模型在熟悉领域能逼近大模型,换个领域就崩——它读出的是基座里已有的东西;
- 合成数据比参数量更关键——数据决定的是「读出哪些知识」;
- 它快——判断所需的计算在执行第一个 token 前就基本完成了。
多模态是更有价值的方向
趋势判断
生产里真正需要判断的东西,大多数带着图。
文本判断点(意图分类、情感、NLI)虽然数量庞大,但很多已经被传统小模型以更低的成本解决了。而带图的判断点——截图、单据、图表、监控画面、网页布局、商品图片——往往没有现成的小模型可用,过去只能靠人工、或者靠昂贵的通用大模型硬扛。
这正是多模态决策模型的机会。而且从数据上看,它已经在领先:Decision Index 的 Vision 子榜最高分 69.8,高于文本主榜的 62.8;已经有 4 个托管端点支持图像输入(Cloudflare Clef / Clef-flash、OpenAI gpt-6-luna-decisions、Perplexity pplx-decider-v1.1-27b),另有 Intern-Decision、decider-2b-vision、JPT 系列等开源选择。
本质是降本增效,终点是业务场景的后训练
观点
先用最强的模型把判断做对,再用它的输出把便宜模型教对。
决策模型的价值不在「更聪明」,而在同一件事更便宜、更快:输入价 $0.02–0.24/百万 token、输出 token 免费、一次前向。省下来的是推理算力、是延迟、是解析与重试的工程成本。
但通用决策模型只解决「起步」问题。真正落地时,你会发现业务判断的长尾极其琐碎:你们公司特有的分类体系、你们自己的合规口径、你们那个行业的术语。这些不会出现在任何公开训练数据里。
于是路径变得清晰:用高阶模型在你自己的数据上跑出结果,再用这些结果去训练或校准一个小决策模型。本质接近蒸馏,只不过蒸出来的不是文本生成能力,而是判断——这正是 arXiv 上那篇《LLM-as-Jev: LLMs Are Already Jev-Style Decision Models — When and How to Fine-Tune Them》在讨论的事情。
所以一个实用的判断标准是:开源且提供完整后训练路径的模型,长期价值更高。只能调 API 的闭源端点适合快速验证,而你要沉淀的是自己的判断能力——那部分资产必须能落在自己的权重上。
按这个标准,目前值得优先研究的开源权重是:Cloudflare Clef / Clef-flash(Apache-2.0,有官方 RL 微调服务)、Perplexity pplx-decider-v1-27b(Apache-2.0,约 49 GiB,可自行 LoRA),以及带完整训练配方的 kev(0.8B 只要 4 GB 显存)、von、laya、decider、bekko(单张 5090 可训)、minojev(只训 0.8M head,笔记本 37 分钟)。多模态场景可以照 jev-smol 的路线:SmolVLM2-500M + 语言塔 LoRA,在 Apple Silicon 上 35 分钟出第一版。
固定槽位头Fixed slots
怎么做:用一个固定宽度的输出层替换原来的语言模型头,每个槽位对应一类问题的答案(例如 256 路选项槽 + 三档分数槽)。
优点:完全不生成 token,延迟最低;输出空间在训练时就确定,不依赖 prompt 渲染。
代价:选项上限写死在权重里;换 schema 必须重训或至少重新校准。
代表实现:OpenThai-SystemOne(256 路槽位替换 LM head)
选项标记Option markers
怎么做:把每个选项的标签当作一段特殊文本拼在候选位置,取该位置上的表示或 logit 作为分数。可以配合顺序无关的注意力掩码。
优点:天然支持可变数量的选项;对选项文本的措辞敏感,可以靠写清楚 criteria 提升准确率。
代价:选项顺序、措辞与数量会影响结果,需要专门做顺序与措辞的鲁棒性测试。
代表实现:Laya(两层 option-marker 打分器,另有 act/escalate 头)、von(option-marker 头 + 顺序无关注意力掩码)、Kev(LoRA + pointer head 读选项 logit)、Decision-1.0(共享候选头,读候选端点与 query 向量)
标签 logit 读出Label logits
怎么做:不新增结构,直接把选项字母/标签 token 放在答案位置做 teacher forcing,取该 token 的 log-prob,再对同题所有选项做 softmax。
优点:实现最轻,只需一次前向;概率直接来自模型分布,配合温度拟合即可校准;任何指令微调过的 LLM 都能改造。
代价:受 tokenizer 影响(多 token 标签要取平均);需要严格保证训练与推理的 prompt 逐字节一致。
代表实现:Bespoke Nimble(直接对允许答案 token 打分,T=2.179)、decider(letter-logit 读出后除以存好的温度,T=1.03–1.94)、jev-lite(在选项字母位置读标签 token)、Open-Jev(由 Yes-minus-No 读出初始化的标量头,T=1.897)、jev-smol(CLEF 式逐选项 logit 打分)
成对打分(交叉编码器)Pair scoring
怎么做:把 (state, question, option) 拼成一个序列整体过一遍编码器,输出一个分数;本质是把 reranker 的结构搬到类型化决策上。
优点:候选与上下文可以互相注意,交互建模最充分;常从检索 reranker 初始化,冷启动快。
代价:每个选项都要过一次前向,选项多时算力线性增长;比前几种慢。
代表实现:open-jev-deberta(三层打分头,池化后的问句与选项表示)、System One scorer(序列分类头逐三元组打分)、GLiNER2.5-Decide(编码器同时做分类、抽 span 与关系)
受限解码Constrained decoding
怎么做:仍然逐 token 生成,但用语法/状态机把输出约束在合法答案集合内,本质是「生成 + 解析」的自动化版本。
优点:改动最小,可以套用在任何现成 LLM 上;能顺带产出简短理由。
代价:没有摆脱自回归,延迟与成本与普通 LLM 同级;概率不是模型分布的直接读出,校准通常最差。
代表实现:djev / razorback16 diffgemma(vLLM 上的 DiffusionGemma 实验路径)、各类把 LLM 变成决策模型的社区适配层
Calibrated Probability
A calibrated probability is a number from a model that means what it says: across all the answers given at 80%, about 80% are right. Jev returns one for every question, a distribution over the options for Choice and Score and a single yes probability for Noul.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Calibrated Probability
Calibrated Probability
A calibrated probability is a number from a model that means what it says: across all the answers given at 80%, about 80% are right. Jev returns one for every question, a distribution over the options for Choice and Score and a single yes probability for Noul.
Full guide: Choice, Score and Noul: the three primitives
The distinction worth holding onto is between a number that ranks and a number that predicts. A softmax output from an ordinary classifier ranks options fine but often runs hot, clustering near 1.0 whether or not the model deserves it. A calibrated probability is meant to be read as a frequency, so 0.7 means seven times in ten.
For a Choice question the probabilities cover the listed options and sum to 1. The docs’ ticket routing example returns {"choice":"returns","confidence":1.0,"probabilities":{"shipping":0.0,"returns":1.0,&qu
Calibration
Calibration is the match between a model's stated probabilities and how often it turns out to be right. A calibrated model that answers 80% across a large batch of questions is correct on about 80% of them. Calibration is about the honesty of the numbers, separate from raw accuracy.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Calibration
Calibration
Calibration is the match between a model's stated probabilities and how often it turns out to be right. A calibrated model that answers 80% across a large batch of questions is correct on about 80% of them. Calibration is about the honesty of the numbers, separate from raw accuracy.
Full guide: RLCD explained: Reinforcement Learning for Calibrated Decisions
Accuracy and calibration come apart. A model can be right 95% of the time and still be badly calibrated if it says 99% every time, and a model that is right only 60% of the time can be perfectly calibrated if it says 60%. The second kind is more useful to software, because a threshold in your code does what you expect.
The docs’ bug severity example shows the shape. A ticket about the export button crashing in Safari gets levels 0 to 2, and the answer comes back with probabilities of 0.0, 0.7 and 0.3, a score of 1.3, and a confidence of 0.54. If the model is calibra
Choice
A Choice is a System One question type for selecting one option from a defined set. The answer names the highest-probability option, includes the full distribution over the options, which sums to 1, and a confidence value from 0 to 1. A Choice question accepts up to 255 options.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Choice
Choice
A Choice is a System One question type for selecting one option from a defined set. The answer names the highest-probability option, includes the full distribution over the options, which sums to 1, and a confidence value from 0 to 1. A Choice question accepts up to 255 options.
Full guide: Choice, Score and Noul: the three primitives
The docs’ example routes a support ticket to a department. With the options shipping, returns and billing, the answer comes back as {"type":"choice","choice":"returns","confidence":1.0,"probabilities":{"shipping":0.0,"returns":1.0,"billing":0.0}}. Your code reads choice to route and reads confidence to decide whether to route at all.
One line of advice in the docs is easy to skip and costly when you do. When your list of options might not cover every real case, add an “other” or “none of the above” option. Without
Composite scoring
Composite scoring is one of TypeSafe's documented patterns, described as "Combine several dimensions of analysis into a single score." You ask a separate question about each dimension, then your own code weights and combines the answers. The model never picks the weights, and the arithmetic stays deterministic.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Composite scoring
Composite scoring
Composite scoring is one of TypeSafe's documented patterns, described as "Combine several dimensions of analysis into a single score." You ask a separate question about each dimension, then your own code weights and combines the answers. The model never picks the weights, and the arithmetic stays deterministic.
Full guide: How to build with System One models
TypeSafe’s build guidance separates the two halves. Ask many independent parallel questions about the same state, then “compose answers with deterministic rules or code-controlled weights”, keeping deterministic work in code. That split matters here, because arithmetic and counting are on jev-1.13’s published list of weak spots. You do not want the model doing the summing.
A lead-scoring example. Ask four Score questions against the same state, each with its own ordered levels: budget fit, urgency, decision-making authority, technical fit. Eac
Confidence Gating
Confidence gating is the practice of branching your code on a model's confidence value as well as on its answer. High confidence runs the action automatically. Low confidence sends the case to a person or a fallback. TypeSafe calls the pattern confidence-gated routing and leaves the thresholds to you.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Confidence Gating
Confidence Gating
Confidence gating is the practice of branching your code on a model's confidence value as well as on its answer. High confidence runs the action automatically. Low confidence sends the case to a person or a fallback. TypeSafe calls the pattern confidence-gated routing and leaves the thresholds to you.
Full guide: How to build with System One models
A classifier that only returns a label forces one behaviour for every prediction. Adding confidence gives a second axis, so the same answer can trigger an automatic action in one case and a review queue in another.
The docs put numbers on it. Anything under 0.5 counts as uncertain and goes to a human. Destructive operations, their example is a funds transfer, need confidence above 0.9 before they run without a confirmation step, while a read-only balance check can proceed at a lower bar. The thresholds scale with what happens if the answer is wrong, and the docs
Confidence
Confidence is a number from 0 to 1 that TypeSafe computes from the probability distribution an answer already carries. A concentrated distribution gives high confidence, a spread one gives low. Choice and Score answers include it. Noul does not, because its single probability already carries the certainty.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Confidence
Confidence
Confidence is a number from 0 to 1 that TypeSafe computes from the probability distribution an answer already carries. A concentrated distribution gives high confidence, a spread one gives low. Choice and Score answers include it. Noul does not, because its single probability already carries the certainty.
Full guide: Choice, Score and Noul: the three primitives
Nothing extra is being measured here. The docs describe confidence as “a statistic computed from the probability distribution the answer already gives you”, collapsing the shape of that distribution into one number. A Choice that puts 1.0 on returns and 0.0 everywhere else comes back with confidence 1.0. A Score split 0.7 and 0.3 across two adjacent levels comes back with confidence 0.54.
TypeSafe’s guidance runs in three bands. High confidence means act automatically. Medium means proceed with caution, asking the user to confirm or flagging the case for review. Low
Decision Model
Decision model is the generic label people reach for when describing an AI model whose output is a typed answer with a probability attached rather than prose. It is not an official term from any vendor. System One model is TypeSafe's name for the same idea, and Jev is the first shipping example.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Decision Model
Decision Model
Decision model is the generic label people reach for when describing an AI model whose output is a typed answer with a probability attached rather than prose. It is not an official term from any vendor. System One model is TypeSafe's name for the same idea, and Jev is the first shipping example.
Full guide: What is a System One (System 1) model?
The phrase gets used two ways. Loosely, it covers any model that hands software a decision rather than a paragraph, which includes encoder classifiers like BERT and DeBERTa, span taggers like GLiNER, and anything wrapped in constrained decoding to force a fixed output shape. More narrowly, people use it for the new category TypeSafe is trying to define, where the typed answer and its probability are what the model was trained to produce in the first place.
There is no standards body here. Treating “decision model” as a description rather than a spec avoids arguing about w
Expected calibration error
Expected calibration error (ECE) is the average gap between a model's stated confidence and how often it is actually right. You bucket predictions into bins by confidence, compare each bin's mean confidence to its accuracy, and average the differences weighted by how many predictions fall in each bin.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Expected calibration error
Expected calibration error
Expected calibration error (ECE) is the average gap between a model's stated confidence and how often it is actually right. You bucket predictions into bins by confidence, compare each bin's mean confidence to its accuracy, and average the differences weighted by how many predictions fall in each bin.
Full guide: RLCD explained: Reinforcement Learning for Calibrated Decisions
A worked version. Take every prediction a model made with confidence between 0.80 and 0.90, where the mean confidence in that bin works out to roughly 0.85. If the model was right on 62% of them, the bin’s gap is about 0.23. Repeat for each bin, weight by bin size, and the total is the ECE. A perfectly calibrated model scores 0, lower is better, and the figure only means something alongside the bin count and the dataset it was measured on.
The measure is standard in the calibration literature. The widely cited ref
Hallucination
Hallucination is when a language model states something false in generated text. A System One model never writes free text, so it cannot invent a fact in prose. It can still be wrong: it can pick the wrong option or return a probability that does not match how often it is right.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Hallucination
Hallucination
Hallucination is when a language model states something false in generated text. A System One model never writes free text, so it cannot invent a fact in prose. It can still be wrong: it can pick the wrong option or return a probability that does not match how often it is right.
Full guide: System One models vs LLMs
The usual failure of a chat model is fluent invention: a citation that does not exist, an API method that was never shipped. Jev returns a typed answer drawn from options you defined, plus a probability distribution over them. There is no prose for a fabricated fact to hide in.
That removes one failure mode. It does not remove error. TypeSafe’s own jaggedness page for jev-1.13 lists nine things the model does badly, including reading dates as text rather than as ordered quantities, and no guarantee that P(A) and P(not A) sum to 1 across related questions. A Choice that comes back wrong at high confidence is
Intent routing
Intent routing is one of TypeSafe's documented patterns: "Classify a user's intent and route to the appropriate handler." A Choice question returns the intent plus a probability for every option, and your code maps the winning option to a handler. The model does not call the handler.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Intent routing
Intent routing
Intent routing is one of TypeSafe's documented patterns: "Classify a user's intent and route to the appropriate handler." A Choice question returns the intent plus a probability for every option, and your code maps the winning option to a handler. The model does not call the handler.
Full guide: How to build with System One models
This is the pattern most people reach for first, and it shows the split between model and code clearly. TypeSafe’s framing: “System One is TypeSafe’s model for building AI-powered software, not agents. It does not generate code or choose its own next action.” The Choice answer is data. Your router is a switch statement.
From the docs, a Choice answer for a support ticket looks like {"type":"choice","choice":"returns","confidence":1.0,"probabilities":{"returns":1.0,"shipping":0.0,"billing&q
Jaggedness
Jaggedness is TypeSafe's word for the published, per-version list of what a model does badly. Each Jev version gets its own page. The one for jev-1.13 documents nine failure modes, so you can design around known weak spots instead of discovering them in production.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Jaggedness
Jaggedness
Jaggedness is TypeSafe's word for the published, per-version list of what a model does badly. Each Jev version gets its own page. The one for jev-1.13 documents nine failure modes, so you can design around known weak spots instead of discovering them in production.
Full guide: What is Jev AI? TypeSafe's decision model explained
Publishing a per-version failure list is unusual, and it helps because the entries are specific. Date and time comparison is one of them: “Jev reads dates as text, not as ordered quantities”, so asking whether one timestamp comes before another is unreliable. Math and numbers is another, with arithmetic and counting called out directly. Adversarial content is a third worth planning around, because “State is data, and jev-1.13 does not treat it as hostile by default”. The fourth is structural: there is no guarantee that P(A) and P(not A) sum to 1 across related questions, so two questions that l
Jev
Jev is the first System One model, announced by TypeSafe AI on 2026-09-15. It answers typed questions about a block of text and returns a choice, a score, or the probability that a yes/no statement is true. Access is a closed managed API, open to anyone since TypeSafe removed the waitlist on 2026-09-20. Weights have not been released.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Jev
Jev
Jev is the first System One model, announced by TypeSafe AI on 2026-09-15. It answers typed questions about a block of text and returns a choice, a score, or the probability that a yes/no statement is true. Access is a closed managed API, open to anyone since TypeSafe removed the waitlist on 2026-09-20. Weights have not been released.
Full guide: What is Jev AI? TypeSafe's decision model explained
A request sends one state, the text to evaluate, plus a set of named questions, and comes back with a typed answer under each name. All the questions in a call run against the same state, so asking four things about one ticket costs one round trip to POST https://api.typesafe.ai/v1/systemone.
The shipping version is jev-1.13.0. The aliases jev-latest and jev-preview both point at it. Input costs $0.042 per million tokens and output is free. A request carries 64k tokens in total, with 32k as the ceiling on the state plus the longest single qu
Non-autoregressive
Non-autoregressive describes a model that does not build its output one token at a time. Autoregressive models predict each token conditioned on the tokens before it. TypeSafe says Jev generates all outputs in a single query instead, though the architecture behind that claim has not been published.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Non-autoregressive
Non-autoregressive
Non-autoregressive describes a model that does not build its output one token at a time. Autoregressive models predict each token conditioned on the tokens before it. TypeSafe says Jev generates all outputs in a single query instead, though the architecture behind that claim has not been published.
Full guide: Jev architecture: parameters, encoder or decoder, and no paper yet
Start with the thing it contrasts against. An autoregressive language model produces text by predicting one token, feeding that token back in, then predicting the next. Each step waits for the one before it, which is why a long answer takes longer to return than a short one.
TypeSafe’s launch post describes Jev’s sampling as “Parallel”, which “Generates all outputs in a single query” versus sequential token generation, and calls it “Incredibly efficient and hardware-aware”. That is the whole public claim. Parameter count, base architectu
Noul
A Noul is a question type in System One AI models such as Jev: it asks a yes/no question and returns one number, the probability that the answer is yes, from 0 to 1. It carries no separate confidence value, because the probability is itself the certainty measure. The docs do not explain the name.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Noul
Noul
A Noul is a question type in System One AI models such as Jev: it asks a yes/no question and returns one number, the probability that the answer is yes, from 0 to 1. It carries no separate confidence value, because the probability is itself the certainty measure. The docs do not explain the name.
Full guide: Choice, Score and Noul: the three primitives
In practice a Noul is a named question in the request with its type set to noul. The docs’ example sends a support message as the state and asks {"type":"noul","instructions":"Is the customer asking for a human agent?"}. The answer comes back as {"type":"noul","noul":0.99}: a 99% probability that the answer is yes. There is no special Noul data type in the response. The value is a plain number, and your code compares it against a threshold like any other float.
Reading the number is direct. A value near 1 is a strong ye
Question
A question is the typed object you attach to state in a System One request. Each one names what you want decided and the shape of the answer. Jev supports three question types: Choice for picking one option, Score for rating against ordered levels, and Noul for a yes/no probability.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Question
Question
A question is the typed object you attach to state in a System One request. Each one names what you want decided and the shape of the answer. Jev supports three question types: Choice for picking one option, Score for rating against ordered levels, and Noul for a yes/no probability.
Full guide: Choice, Score and Noul: the three primitives
Questions are keyed in a map, so each answer comes back under the name you gave it. Choice selects one option from a defined set and returns the winning option, the full probability distribution and a confidence value, with up to 255 options allowed. Score evaluates content against ordered, descriptive levels, minimum 2 and maximum 10, and returns a probability-weighted mean of the level numbers. Noul returns a single number between 0 and 1, the probability that a yes/no statement is true, and no separate confidence value.
The docs advise asking “the most explicit, narrow, specific, atomic ques
RLCD
RLCD stands for Reinforcement Learning for Calibrated Decisions, the training method TypeSafe AI says it used for Jev. Per the launch post it optimizes for answers with epistemically honest probabilities on System One tasks, rather than for text a human rater prefers or output a program can check.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
RLCD
RLCD
RLCD stands for Reinforcement Learning for Calibrated Decisions, the training method TypeSafe AI says it used for Jev. Per the launch post it optimizes for answers with epistemically honest probabilities on System One tasks, rather than for text a human rater prefers or output a program can check.
Full guide: RLCD explained: Reinforcement Learning for Calibrated Decisions
The name sits alongside two older acronyms. RLHF, reinforcement learning from human feedback, trains a model toward responses human raters prefer. RLVR, reinforcement learning from verifiable rewards, trains it toward outputs a program can check as correct. RLCD swaps the target again: the reward is tied to whether the stated probability matches how often the answer turns out to be right.
That target is what calibration means. A model trained this way should be right on roughly 80% of the answers it labels 80%, and wrong on roughly 20% of them, which is what makes a th
Score
A Score is a System One question type that rates content against ordered, descriptive levels, minimum 2 and maximum 10. The returned score is the probability-weighted mean of the level numbers, so it can land between levels. The answer also carries a legend, the level probabilities and a confidence value.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Score
Score
A Score is a System One question type that rates content against ordered, descriptive levels, minimum 2 and maximum 10. The returned score is the probability-weighted mean of the level numbers, so it can land between levels. The answer also carries a legend, the level probabilities and a confidence value.
Full guide: Choice, Score and Noul: the three primitives
Levels are written as short descriptions, which is what lets the model place content without a separate rubric hidden in your prompt. The docs’ bug severity example uses three: “Cosmetic; no impact to functionality”, “Broken or degraded feature, but workaround exists”, and “Blocking issue; no workaround exists”.
Given the state “The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari”, the answer is a score of 1.3 with probabilities of 0.0, 0.7 and 0.3 on levels 0, 1 and 2, at confidence 0.54. The arithmetic is plain
Speculative fan-out
Speculative fan-out is one of TypeSafe's documented patterns: "Send many questions in a single call, including speculative ones, and let your code decide what's relevant." Because questions are evaluated in parallel, extra questions typically cost no extra latency, so you ask ahead instead of making a second round trip.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Speculative fan-out
Speculative fan-out
Speculative fan-out is one of TypeSafe's documented patterns: "Send many questions in a single call, including speculative ones, and let your code decide what's relevant." Because questions are evaluated in parallel, extra questions typically cost no extra latency, so you ask ahead instead of making a second round trip.
Full guide: How to build with System One models
The docs state the mechanism directly: “All questions are evaluated in parallel, so adding more questions to a call typically doesn’t add any latency to the response.”
That flips the usual cost model. With a chat model you ask for the minimum, because every extra field is more tokens to generate and more time to wait. Here the marginal question is close to free in latency terms, so you ask everything any downstream branch might plausibly need, then throw away what the branch you took did not use.
Concrete shape. One call on a
State
State is TypeSafe's word for the content you hand a System One model: the ticket, document, transcript or JSON object the questions are asked about. It is passed as text, and the state plus the longest question must fit within 32k tokens on Jev.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
State
State
State is TypeSafe's word for the content you hand a System One model: the ticket, document, transcript or JSON object the questions are asked about. It is passed as text, and the state plus the longest question must fit within 32k tokens on Jev.
Full guide: How to build with System One models
State is the data half of a System One request. The other half is the set of typed questions you attach to it. Jev accepts text only, passed as a string or as structured text (a JSON object or an array of text values). No image, audio or video input. Total context is 64k tokens per request, and state plus the longest single question is capped at 32k tokens.
TypeSafe’s build guidance is blunt about what belongs in there: “Include only the context relevant to the current questions. This helps the model avoid distractions and context rot.” That advice has teeth. The published jaggedness page for jev-1.13 lists large state full of irrelevant deta
Structured outputs
Structured outputs is the existing technique for making a language model emit valid JSON: constrain decoding to a schema so every generated token keeps the output parseable. OpenAI and others ship it in their APIs, and libraries such as Outlines and Instructor do the same over open models.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Structured outputs
Structured outputs
Structured outputs is the existing technique for making a language model emit valid JSON: constrain decoding to a schema so every generated token keeps the output parseable. OpenAI and others ship it in their APIs, and libraries such as Outlines and Instructor do the same over open models.
Full guide: System One models vs LLMs
Mechanically, schema-constrained decoding masks the logits at each step so only tokens the grammar allows can be sampled. The model still generates a token at a time. A JSON object with six fields costs six fields’ worth of sequential decoding, and any confidence figure has to be reconstructed from logprobs afterwards.
A System One model differs at exactly that step. Nothing is generated and then constrained, because the option set you passed in is the answer space, and a calibrated probability over those options comes back as part of the answer. TypeSafe’s public claim is that Jev “Gen
System One Model
A System One model is a class of AI models built to make fast, structured decisions that software can use directly. It returns typed decisions and probabilities instead of generated text, so it does not write replies, produce code, or explain its reasoning. Jev by TypeSafe AI is the first one.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
System One Model
System One Model
A System One model is a class of AI models built to make fast, structured decisions that software can use directly. It returns typed decisions and probabilities instead of generated text, so it does not write replies, produce code, or explain its reasoning. Jev by TypeSafe AI is the first one.
Full guide: What is a System One (System 1) model?
The category name comes from Daniel Kahneman’s Thinking, Fast and Slow, which splits thinking into fast, intuitive System 1 and slow, deliberate System 2 reasoning. A System One model takes the fast half: a judgment your code needs in milliseconds, returned in a shape your code can branch on without parsing prose. TypeSafe’s docs define three question types for it, Choice, Score and Noul, and the model never picks the next action, so your code keeps the control flow and every side effect.
This entry is the short definition. The full explanation, with the three question type
System Two Model
System Two model is a contrast term, not a product. No vendor ships a model branded this way. The phrase points to the slow, deliberate half of Daniel Kahneman's split in Thinking, Fast and Slow, which in practice means a reasoning LLM that works through text before it answers.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
System Two Model
System Two Model
System Two model is a contrast term, not a product. No vendor ships a model branded this way. The phrase points to the slow, deliberate half of Daniel Kahneman's split in Thinking, Fast and Slow, which in practice means a reasoning LLM that works through text before it answers.
Full guide: System One models vs LLMs
The term only exists because TypeSafe named its own category after Kahneman’s System 1. Its launch post cites “fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning” as the source of the name. Once one half has a product label, people reach for the other half to describe everything that is not a System One model: chat models and reasoning models, anything that works in tokens and hands you prose.
Nobody sells a “System Two model”. If you see the phrase in a comparison table, read it as shorthand for a reasoning LLM.
The practical difference is time budget. TypeSafe claims Jev a
Thinking, Fast and Slow
Thinking, Fast and Slow is Daniel Kahneman's 2011 book (Farrar, Straus and Giroux) describing two modes of human thought: a fast, automatic System 1 and a slow, deliberate System 2. TypeSafe borrowed the framing for the name "System One model", the category Jev launched under in September 2026.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Thinking, Fast and Slow
Thinking, Fast and Slow
Thinking, Fast and Slow is Daniel Kahneman's 2011 book (Farrar, Straus and Giroux) describing two modes of human thought: a fast, automatic System 1 and a slow, deliberate System 2. TypeSafe borrowed the framing for the name "System One model", the category Jev launched under in September 2026.
Full guide: What is a System One (System 1) model?
The launch post makes the borrowing explicit, contrasting “fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning” and placing the new model class on the fast side: quick typed decisions that software consumes, with no written reasoning attached.
Kahneman’s subject was people. System 1 in the book describes human cognition, its speed and its characteristic biases, and makes no claim about how any machine is built. Using the name for a class of models is an analogy about the job being done, not evidence about architecture. What
Typed output
Typed output is a model answer that arrives already shaped as a value your code can use, such as an enum member or a probability between 0 and 1. There is no parsing step and no schema validation on a text blob. A System One model returns typed output directly rather than describing an answer in prose.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Typed output
Typed output
Typed output is a model answer that arrives already shaped as a value your code can use, such as an enum member or a probability between 0 and 1. There is no parsing step and no schema validation on a text blob. A System One model returns typed output directly rather than describing an answer in prose.
Full guide: System One models vs LLMs
TypeSafe describes System One models as returning “typed answers and probabilities rather than generated text”, and says they “do not write replies, produce code, or generate explanations of their reasoning.”
What comes back from a Choice question is concrete. The docs show {"type":"choice","choice":"returns","confidence":1.0,"probabilities":{"returns":1.0,"shipping":0.0,"billing":0.0}}. In the JavaScript SDK you read it as response.answers.category.choice, and the value is one of the keys you sup
TypeSafe AI
TypeSafe AI is the San Francisco company that coined the term System One model and shipped the first one, Jev, on 2026-09-15. It was founded by Diogo Almeida (CEO), Erik Gafni (CTO) and Sasha Sheng (COO), and raised a $40M seed round led by DCVC.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
TypeSafe AI
TypeSafe AI
TypeSafe AI is the San Francisco company that coined the term System One model and shipped the first one, Jev, on 2026-09-15. It was founded by Diogo Almeida (CEO), Erik Gafni (CTO) and Sasha Sheng (COO), and raised a $40M seed round led by DCVC.
Full guide: What is Jev AI? TypeSafe's decision model explained
The company’s pitch starts from a question Almeida puts at the top of the launch post: “Models have been superhuman at chat for years, so where is all the automation?” Its answer is a model that returns a typed decision your code can branch on, with a probability attached, instead of a paragraph you have to parse.
TypeSafe’s own team page says Almeida co-invented RLHF and InstructGPT and previously worked at Google Brain, Gafni is a repeat founder with production AI experience, and Sheng was a research engineer at Meta’s FAIR. That background matters mainly because the training method behind Jev, RLCD, is describe
Zero-shot classifier
A zero-shot classifier assigns text to labels it was never trained on, using the label descriptions you supply at call time rather than labelled examples. Jev's Choice question works this way: you define the categories in the request, and the model returns a probability distribution over them.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu
Site navigation
ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example
Close
Home
Glossary
Zero-shot classifier
Zero-shot classifier
A zero-shot classifier assigns text to labels it was never trained on, using the label descriptions you supply at call time rather than labelled examples. Jev's Choice question works this way: you define the categories in the request, and the model returns a probability distribution over them.
Full guide: Is Jev just a zero-shot classifier?
The term predates Jev by years. Encoder classifiers such as BERT and DeBERTa, and span models such as GLiNER, have been doing classification against labels supplied at inference time for a while, which is why the comparison came up straight after the launch on 2026-09-15.
agentpedia.codes reports that a Hacker News commenter called Jev “basically a zero-shot classifier” and that Diogo Almeida replied “exactly right!”. That exchange could not be confirmed in the Hacker News thread itself, so treat it as agentpedia.codes’ account rather than a direct citation from HN
System One 模型(System 1)
返回类型化决策与校准概率、而不是生成文本的一类 AI 模型。
展开
名字借自双过程理论:System 1 是快速直觉,System 2 是慢速推理。System One 模型承担前者——在流程里做快速、结构化、可直接被代码消费的判定。第一个是 TypeSafe AI 的 Jev(2026-09-15)。
RLCD
Reinforcement Learning for Calibrated Decisions,TypeSafe 用于 Jev 的训练方法。
展开
公开信息只到博客级别:给相邻的有序选项部分分、奖励完全正确的整条记录输出、并加一个 reference penalty 防止分布漂移。具体损失与数据未公开,也没有公开实现。
Brier 分数
衡量概率预测准确度的严格评分规则,越低越好。
展开
对二分类是 (p - y)² 的均值。S1MB 的 Noul 任务用它作为主指标,并用 balanced accuracy 做基线校正后折算成 0–100 的任务分。
ECE(期望校准误差)
把预测按置信度分桶,比较每桶的平均置信度与实际准确率。
展开
ECE 越小说明「说 90% 就是 90%」。决策模型的价值很大程度依赖这个性质,因为代码要拿概率去比阈值。Decision Index 对每个模型报告 ECE、Brier 与 over95(置信度 >95% 的错误率)。
per-option logit 打分
不生成文本,而是把每个选项的标签放在答案位置做 teacher forcing,取平均 token log-prob 再 softmax。
展开
这是当前开源决策模型最主流的实现方式(jev-smol、decider、JevK5/SemIf、Intern-Decision 等)。相比「生成 JSON 再解析」,它更快、更稳、也更好校准,因为概率直接来自模型分布。
joint schema head
在冻结 backbone 的隐状态上加一个小 transformer head,联合给所有问题的所有选项打分。
展开
Cloudflare Clef 的架构:backbone 只做一次 prefill,head 负责把证据路由到每个问题、并让字段之间互相关注,最后一次性输出所有选项的 logit。
confidence 与概率的区别
probability 是每个选项的分布,confidence 是模型对「这次选择」的把握。
展开
Choice 与 Score 会同时返回两者;Noul 在 Jev 上只返回概率、不返回独立 confidence。生产系统通常对 confidence 设阈值,低置信度转人工或升级到更强的模型。
基线校正分(S1MB Task Avg)
不是准确率,而是「相对忽略输入的基线补上了多少差距」。
展开
50 分意味着补上了一半差距。Noul 用 balanced accuracy 把 50% 映射到 0、100% 映射到 100;Choice 取「均匀随机」与「固定作答」中较优者为基线;Score 用 MAE 对常数预测做校正。
Borda Score
把每个基准上的相对名次折算成分数的排序方法。
展开
每个基准第一名 100 分、最后一名 0 分,中间等距;再对所有基准等权平均。它衡量「稳定地名列前茅」,会随参赛模型名单变化而变化,不能与 Task Avg 混用。
决策原语(primitive)
Choice / Noul / Score 三种问法,对应分类、判定与打分。
展开
同一份 state 可以挂多个问题、不同类型混用,一次前向返回全部答案。OpenAI 把它叫 predicate / choice / score,Perplexity 与 S1MB 直接叫 noul / choice / score。