decision.host
首页 / 模型 / kev-0.5b

kev-0.5b

Jared Palmer ChoiceNoulScore 开源权重 4.2B(活跃 358M) curateddiopenrouters1mbsom

覆盖 0.5B / 0.6B / 0.8B / 4B / 8B / 9B / 27B 的模型家族,Apache-2.0,TypeSafe 兼容 API,4B 同时在 OpenRouter 上以 $0.042/百万提供。如果你要微调自己的决策模型,这是入门首选。

S1MB Task Avg
10.1
第 78 名 · 覆盖 137/137(100%)

模型信息

厂商
Jared Palmer
类别
开源权重
参数规模
4.2B(活跃 358M)
基座模型
Qwen3.5 / Qwen3.8 系列
权重
可下载(开源)
许可证
Apache-2.0 可商用
输入模态
文本
决策原语
ChoiceNoulScore
上下文
8,192 token
输入价格
$0.042/M tok
延迟
41.5-145.2 ms
可微调
可以
微调方法
文档最完整的开源决策模型:koev.train --init_from <ckpt>;LoRA + pointer head;逐 checkpoint 拟合温度(0.8B T=2.35 / 9B T=2.30)
微调硬件
0.8B bf16 仅需 4 GB 显存;4B 走 Modal 约 $1/H100
训练方式
full fine-tune
OpenRouter 模型 ID
jaredpalmer/kev-4b
OpenRouter 计费
输入 $0.042/M tok · 输出 免费
上架状态
Generally available

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
laya / ag news Choice 89.5 94.0 75.4 来源 ↗
dbpedia Choice 80.9 98.9 83.5 来源 ↗
banking77 Choice 80.0 93.7 54.6 来源 ↗
s1mb generalization contextual choice Choice 76.8 100.0 82.6 来源 ↗
open jev / browser control v1 Choice 73.9 100.0 67.4 来源 ↗
open jev / customer control v1 Choice 73.9 82.6 64.3 来源 ↗
open jev / phone extraction control v1 Noul 63.6 100.0 50.4 来源 ↗
massive Choice 59.4 96.7 64.5 来源 ↗
mtop Choice 56.6 86.9 58.6 来源 ↗
open jev / mailroom control v1 Choice 50.0 100.0 82.9 来源 ↗
open jev / amount extraction control v1 Choice 49.0 100.0 55.0 来源 ↗
hwu64 Choice 48.9 92.5 56.8 来源 ↗
sdoh nli Noul 48.0 92.0 60.8 来源 ↗
miqa Choice 47.7 95.5 61.9 来源 ↗
s1mb generalization diverse choice Choice 47.1 100.0 69.5 来源 ↗
clinc Choice 45.5 99.3 57.3 来源 ↗
scicite Choice 45.2 71.0 44.0 来源 ↗
open jev / painting geometry v1 Choice 42.1 100.0 46.2 来源 ↗
canttalk Noul 42.0 86.0 43.9 来源 ↗
winowhy Noul 40.0 62.0 18.8 来源 ↗
arc Choice 33.8 98.5 64.4 来源 ↗
go emotions Noul 28.9 61.3 38.3 来源 ↗
scitail Noul 26.0 98.0 53.0 来源 ↗
aegis2 Noul 22.8 62.3 32.8 来源 ↗
followir robust04 Noul 22.0 90.9 46.9 来源 ↗
babi nli Noul 22.0 84.0 37.0 来源 ↗
open jev / mailroom control v1 Noul 21.1 100.0 62.5 来源 ↗
followir news21 Noul 20.7 42.7 14.9 来源 ↗
open jev / customer control v1 Score 20.5 99.3 36.7 来源 ↗
ethos Noul 20.0 76.0 43.7 来源 ↗
open jev / email selection control v1 Noul 19.2 100.0 40.6 来源 ↗
open jev / reasoning control v1 Noul 16.2 100.0 25.3 来源 ↗
laya / typed decisions Noul 16.1 47.8 24.4 来源 ↗
hatecheck Noul 16.0 98.0 47.5 来源 ↗
civil comments Noul 14.8 36.7 21.4 来源 ↗
impli Noul 14.0 88.0 46.7 来源 ↗
qasper Noul 14.0 88.0 41.6 来源 ↗
s1mb generalization contextual score Score 13.9 92.7 42.1 来源 ↗
gretel pii Choice 13.6 94.5 43.1 来源 ↗
few nerd Choice 13.6 86.4 43.9 来源 ↗
openbookqa Choice 12.5 97.2 53.7 来源 ↗
creak Noul 10.0 88.0 38.8 来源 ↗
open jev / amount extraction control v1 Noul 9.4 100.0 33.3 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 9.3 70.8 28.4 来源 ↗
ethics Noul 9.1 68.6 23.9 来源 ↗
laya / typed decisions Choice 8.3 34.5 18.9 来源 ↗
s1mb generalization contextual noul Noul 8.0 98.0 62.0 来源 ↗
open jev / workflow controls v1 / customer service Noul 8.0 100.0 41.6 来源 ↗
open jev / ir control v1 Noul 7.1 100.0 47.3 来源 ↗
lexcomp Noul 6.0 52.0 22.5 来源 ↗
s1mb generalization diverse noul Noul 6.0 100.0 59.2 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 4.2 52.1 29.0 来源 ↗
paws Noul 4.0 88.0 45.7 来源 ↗
sgd Choice 3.4 96.6 24.4 来源 ↗
laya / typed decisions Score 3.3 63.5 26.5 来源 ↗
quartz Choice 2.4 88.1 49.1 来源 ↗
arct Choice 2.1 89.4 39.4 来源 ↗
laya / enron spam Noul 2.0 100.0 54.6 来源 ↗
bbq Choice 1.7 98.3 53.8 来源 ↗
robust lr Choice 1.7 66.7 18.8 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 1.7 47.0 26.6 来源 ↗
fol nli Choice 1.6 54.1 10.4 来源 ↗
snli Choice 1.6 98.4 46.0 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 0.7 54.0 28.6 来源 ↗
aqua rat Choice 0.0 77.5 10.7 来源 ↗
argument quality Choice 0.0 95.9 37.4 来源 ↗
argument quality Noul 0.0 23.9 5.4 来源 ↗
boardgameqa Choice 0.0 62.5 15.6 来源 ↗
cladder Noul 0.0 94.0 17.7 来源 ↗
contract nli Choice 0.0 76.0 31.8 来源 ↗
corr2cause Choice 0.0 91.7 3.8 来源 ↗
crows pairs Choice 0.0 70.8 30.4 来源 ↗
defeasible nli Choice 0.0 80.4 32.8 来源 ↗
esci Choice 0.0 46.6 13.0 来源 ↗
ethics Choice 0.0 83.3 20.9 来源 ↗
followir core17 Noul 0.0 46.6 21.2 来源 ↗
gsm8k Choice 0.0 78.6 23.5 来源 ↗
hans Noul 0.0 100.0 43.3 来源 ↗
hh rlhf Choice 0.0 23.1 3.1 来源 ↗
logical entailment Noul 0.0 72.0 12.2 来源 ↗
lonli Choice 0.0 86.9 39.2 来源 ↗
nlsat Noul 0.0 22.0 2.1 来源 ↗
open jev / citation control v1 Choice 0.0 100.0 40.9 来源 ↗
open jev / context retention control v1 Noul 0.0 100.0 19.0 来源 ↗
open jev / customer control v1 Noul 0.0 100.0 58.0 来源 ↗
open jev / drone control v1 Choice 0.0 100.0 4.8 来源 ↗
open jev / drone control v1 Noul 0.0 100.0 50.2 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
open jev / email selection control v1 Choice 0.0 100.0 23.4 来源 ↗
open jev / entity alignment control v1 Noul 0.0 100.0 35.9 来源 ↗
open jev / entity alignment control v1 Score 0.0 100.0 8.1 来源 ↗
open jev / ir control v1 Choice 0.0 88.0 44.8 来源 ↗
open jev / ir control v1 Score 0.0 98.6 10.1 来源 ↗
open jev / painting geometry v1 Noul 0.0 100.0 27.2 来源 ↗
open jev / painting geometry v1 Score 0.0 100.0 15.2 来源 ↗
open jev / phone extraction control v1 Choice 0.0 100.0 32.4 来源 ↗
open jev / reasoning control v1 Choice 0.0 95.0 28.4 来源 ↗
open jev / reasoning control v1 Score 0.0 90.7 27.7 来源 ↗
open jev / silent failure control v1 Noul 0.0 100.0 38.8 来源 ↗
open jev / snake v1 Choice 0.0 83.5 8.1 来源 ↗
open jev / snake v1 Noul 0.0 100.0 9.0 来源 ↗
open jev / sponsor segment control v1 Choice 0.0 100.0 52.2 来源 ↗
open jev / tic tac toe v1 Choice 0.0 48.5 3.5 来源 ↗
open jev / vizdoom basic v1 Noul 0.0 100.0 39.0 来源 ↗
open jev / vizdoom basic v1 Score 0.0 100.0 29.6 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 0.0 100.0 27.8 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 0.0 100.0 34.6 来源 ↗
open jev / workflow controls v1 / security incidents Noul 0.0 100.0 37.8 来源 ↗
patent similarity Score 0.0 48.3 17.7 来源 ↗
plane Noul 0.0 100.0 11.1 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗
ruletaker Noul 0.0 96.0 27.7 来源 ↗
s1mb generalization diverse score Score 0.0 95.3 42.8 来源 ↗
scone Choice 0.0 92.0 34.2 来源 ↗
spartqa Choice 0.0 38.7 4.4 来源 ↗
stepgame Choice 0.0 82.1 17.8 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 0.0 48.2 15.6 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 0.0 36.0 7.7 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 0.0 44.9 17.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 0.0 46.7 15.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 0.0 63.6 13.9 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 0.0 62.5 37.3 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 0.0 52.6 17.8 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 0.0 46.2 25.1 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 0.0 38.6 10.7 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 0.0 67.2 39.5 来源 ↗
synthetic relevance / nanobeir / nanonq Score 0.0 49.4 9.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 0.0 41.7 17.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 0.0 36.7 7.2 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 0.0 44.3 20.6 来源 ↗
synthetic relevance / nanocoir Noul 0.0 73.0 33.6 来源 ↗
synthetic relevance / nanocoir Score 0.0 44.3 4.3 来源 ↗
temporal nli Choice 0.0 80.7 18.3 来源 ↗
tracie Noul 0.0 34.0 8.7 来源 ↗
ud ewt Choice 0.0 56.7 13.2 来源 ↗
winobias Choice 0.0 91.3 30.1 来源 ↗
wiqa Choice 0.0 58.7 19.3 来源 ↗

同族条目(9)

型号参数Decision Index S1MBVision榜单
kev-0.5b 当前 4.2B(活跃 358M) — 10.1 — s1mb
Kev 27B 28B 58.8 — — di
jaredpalmer/kev-9b 7.9B(活跃 6.9B) 43.3 42.4 — s1mb, di
Kev 4B r10 4.7B 39.5 — — di
Kev 0.8B r15 870M 13.7 — — di
kev-0.6b 597M(活跃 441M) — 14.1 — s1mb
jaredpalmer/kev-0.8b 753M(活跃 499M) — 18.7 — s1mb
jaredpalmer/kev-4b 4.2B(活跃 3.6B) — 42.3 — s1mb
kev-8b 7.6B(活跃 6.9B) — 36.0 — s1mb

来自 systemonemodels.org 的详细介绍

What Kev is

Kev is a family of open-source System One models published by Jared Palmer. It reads a state and answers typed Choice, Score and Noul questions about it with probabilities, the same three shapes Jev uses. You download the weights from Hugging Face and run them on CUDA, ROCm or Apple Silicon through MLX, or call Kev-4B on OpenRouter. The 27B model needs an 80 GB data-centre GPU and has no Mac path. The server exposes POST /v1/systemone with the same request and response shape as TypeSafe’s API, so the typesafe-sdk Python client works once you point base_url at your own machine. The repo also ships a Modal script that deploys Kev-4B to an L40S behind a bearer key, and a Hugging Face Space runs Kev-4B and Kev-0.8B in the browser. OpenRouter added Kev-4B on 25 September 2026, hosted by SiliconFlow in fp8. It costs $0.042 per million input tokens with free output and an 8,192-token context, the same price as Jev. OpenRouter says it takes the same request as Jev, so switching means changing the model field to jaredpalmer/kev-4b. The 0.8B, 9B and 27B models are not hosted there.

How it is built

Each checkpoint is a rank-16 LoRA adapter plus a small pointer head on a Qwen base: Qwen3.5 at 0.8B, 4B and 9B, and Qwen3.8 at 27B. The head scores each option against a decision token at the end of its question, and a softmax turns those scores into probabilities. Questions share the state but cannot see each other. Training is plain cross-entropy on 10,000 examples from ten public datasets, plus 896 generated policy examples and 1,680 from generated rule structures. The README says no Jev outputs were used. Kev-27B trains for one epoch on Kev-9B’s data, plus 1,400 records where the question sits among 1k to 6k tokens of unrelated text, and it uses soft targets where the right answer is genuinely ambiguous. Each checkpoint ships with a temperature fitted on in-distribution data, which changes the probabilities but never the answer. The 0.8B, 4B and 9B models got a short second training pass on generated examples on 2026-09-21, with the earlier weights kept at revision v7-base. Kev-4B was updated again on 2026-09-24 after one epoch on consumer-finance complaints. Kev-27B differs from the others in its base. The smaller models start from Qwen base checkpoints. Kev-27B starts from Qwen3.8-27B, Qwen’s post-trained release, and the README says its training data is unknown. Comparisons with Jev or with the smaller Kevs therefore do not test the method alone.

What its own evals show

These are the author’s numbers from the repo’s own benchmark, which also runs the Jev comparison. On datasets Kev was not trained on, Kev-9B scores 0.822 on the development set against Jev’s 0.857, and 0.852 on a test set Jev has not been run on. Kev-27B scores 0.848 on the development set, within a point of Jev, and 0.896 on the test set. On held-out examples from its training sources Kev-9B scores 0.872 against Jev’s 0.845, and Kev-27B 0.866. The README says this is not a controlled comparison, because Jev’s training data is unknown, and Kev-27B’s post-trained base adds a second unknown. The first independent test is the Decision Index 0.2.1 by multimodalart, updated 28 September 2026. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Jev scores 57.91, Kev-9B 38.48, Kev-4B 34.64 and Kev-0.8B 14.60. Kev-9B ranks 26th of the 70 entries, 19.4 points behind Jev. Their expected calibration error, the average gap between stated confidence and actual accuracy, is 0.138 for Kev-9B, 0.176 for Kev-4B and 0.074 for Kev-0.8B, against Jev’s 0.074. Kev-27B was not in the index on 30 September. Speed is self-reported too. Kev-4B on an L40S takes 41.5ms to 145ms of model time for a new state, depending on how many questions you ask and how long the state is. Kev-27B takes 46.5ms to 178ms on a B200 and 75ms to 278ms on an H100. The README’s sample request on an Apple M5 in bf16 came back in 495ms.

What it is not for

Knowledge questions are the widest gap: Kev-9B scores 0.74 on MMLU and Kev-27B 0.84, where Jev scores 0.90. Changing the order of options can change an answer. Training covered at most 384 state tokens, so long documents fall outside it even though the server accepts a 65,536-token state. The server handles one request at a time, and the Qwen3.5 models run slowly on a Mac. Kev-27B needs an 80 GB GPU, with 55 GB of weights and about 66 GB in use once the serving buffers are counted, and it has no Mac path.

Specifications

Question types Choice Score Noul Max Choice options255 Score levelsUp to 255 Questions per callNot documented Total context8,192 tokens State budgetNot documented Rate limitNone when self-hosted, where the server handles one request at a time. OpenRouter does not publish a rate limit for Kev-4B. EndpointPOST /v1/systemone on your own server, or Kev-4B through OpenRouter SDKsPython: typesafe-sdk The README says the server accepts a state of up to 65,536 tokens plus 8,192 for each question, but training used at most 384 state tokens and 1,024 for the state plus one question, so longer inputs fall outside what training covered. Kev-27B holds up better on long text: 0.833 on questions buried in 1k to 6k tokens of unrelated text, against 0.556 for Kev-9B. OpenRouter lists an 8,192-token context for Kev-4B. Choice takes 1 to 255 options and Score 1 to 255 levels. A request can carry any number of questions; the server runs them a 16,384-token row at a time, so memory does not grow with the question count.

Versions

jaredpalmer/kev-0.8b, 21 Sep 2026, LoRA adapter and pointer head on Qwen3.5-0.8B-Base. Weights updated on 2026-09-21 with a second training pass on generated examples; the previous weights are at revision v7-base. Release notes jaredpalmer/kev-4b, 24 Sep 2026, The recommended starting point, on Qwen3.5-4B-Base. Updated on 2026-09-21 like the others, then again on 2026-09-24 with one epoch on 5,219 consumer-finance complaint narratives; the previous weights are at revision night2-du-release. OpenRouter serves this build as kev-4b-20260924. Release notes jaredpalmer/kev-9b, 21 Sep 2026, The most accurate Kev on a Qwen3.5 base. Updated on 2026-09-21; the previous weights are at revision v7-base. Release notes jaredpalmer/kev-27b, 24 Sep 2026, The most accurate Kev, a LoRA adapter and pointer head on Qwen3.8-27B, Qwen's post-trained release rather than a base checkpoint. Apache 2.0. It needs an 80 GB GPU (55 GB of weights in bf16, about 66 GB with serving buffers) and has no Mac path. Release notes

Use cases

What people use Kev for, one page per pattern. Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket. Choice Score Noul Workflow controlIntermediate

Confidence-gated actions with System One models

Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars. Choice Score

Examples built with Kev

The most-starred and most-viewed entries in the directory. Browse all examples.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

Cloudflare's launch post for Clef and Clef-flash, two decision models on Workers AI that accept the Jev request format and open weights under Apache 2.0. It covers the architecture, Cloudflare's own benchmark tables against Jev, Kev and Laya, and a new reinforcement learning service. Michelle Chen (@michellechen) 140k viewsOpen Introducing Clef: our open-source decision models, and new RL fine-tuning platform on blog.cloudflare.com Article Support inbox triage with System One models

Decision Model Leaderboard

Cloudflare's live leaderboard for decision models on the Decision Index suite, with Clef, Clef-flash, Jev, Kev-9B and Laya plotted by score and latency. Cloudflare built and ran it, so treat the ranking as vendor-run. Cloudflare (@ritakozlov) 41k viewsOpen Decision Model Leaderboard on clef-evals.workers-ai-mle.workers.dev Article