decision.host

决策模型

决策模型(Decision Model,也叫 System One 模型或类型化概率决策模型)是一类只做判断、不写文章的模型。给它一段内容(state)和几个类型化的问题答案,它直接返回答案概率。它不生成解释、不做推理链、不输出自由文本。这个品类的起点是 2026 年 9 月 15 日 TypeSafe AI 发布的 Jev 1.13,此后进入了高速发展期。本站会实时更新收录最新的模型、评估、论文等等相关进展。

193模型
246评估集
17,791得分记录
1,192相关论文
82开源权重
28多模态

榜单速览

口径不同,分数不可跨榜换算

S1MB Task Avg(英文文本)

完整 →
1 OpenJev-27B 62.6
2 AutoJev-27B 60.8
3 Eikos 27B 59.9
4 Jev 1.13 59.6

Decision Index Vision(多模态)

完整 →
1 JEV-27B-VL 69.8
5 JPT-9B 61.9
2026-10-06 方法与训练 相关度 43.5 cs.CL
Shuyu Gan、Young-Jun Lee、Dongyeop Kang
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.
决策模型System 1 双过程类型化决策决策头 / 读出严格评分规则
2026-10-06 方法与训练 相关度 28 cs.CL
Ran Li、Lei Chen
Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label menu.Code and data are available at https://github.com/rlisml/jev-cascade.
决策模型Jev / TypeSafe单次前向 / prefill 读出LLM 校准
2026-10-06 系统与工程 相关度 10 cs.AI
Ankit Sonthalia、Haritz Puerto、Alexander Rubinstein 等 5 人
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model's zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a "system one" model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents' ability to invest limited resources in reusable solutions for large, repetitive workloads.
System One 模型Jev / TypeSafe
2026-10-06 系统与工程 相关度 8 cs.MA
Zihan Zhou、Xinzhe Hu、Hanxu Yang 等 5 人
Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
System One 模型

最新动态

全部 →
2026-10-06
Decision Index 0.3 发布:Perplexity Decider v1.1 登顶,Jev 掉到第 3
新版权重为 0.2×public + 0.5×same_skills + 0.3×new_domains。Perplexity Decider v1.1 (27B) 以 62.8 分第一,Fastino GLiDE 60.2 第二,Jev 1.13 60.1 第三。同版还发布了 Vision 子榜,JEV-27B-VL 以 69.8 分领先。
2026-10-06
Intern-Decision 0.8B/2B/4B 发布:微调语言主干、冻结视觉塔
上海 AI Lab 系的开源多模态决策模型,每字段一个 <decision> token,单张 RTX 4090 约 33 ms/query,训练代码一并开源。
2026-10-01
Cloudflare 开源 Clef / Clef-flash,并推出 RL 微调平台
Clef 27B(Qwen3.8-27B 底座)、Clef-flash 9B(Qwen3.5-9B),Apache-2.0,带视觉编码器、64K 上下文,一轮前向并行给所有选项打分。训练用 label-smoothed CE + Brier loss 做校准,再加 RLCD。这是第一个既有开源权重、又有官方微调服务的决策模型。
2026-10-01
OpenAI 推出 Decisions API(public beta),也是多模态的
端点 POST /v1/decisions,只有 gpt-6-luna 一个模型,支持 text + image 输入,三种问法 predicate / choice / score。输入 $0.10/百万 token,输出免费。支持 ZDR 与 HIPAA,区域限美国与欧洲。

数据来源

每次更新都会记录抓取时间与校验值