decision.host
首页 / 模型 / togethercomputer/Tev1-4B-experimental

togethercomputer/Tev1-4B-experimental

Together AI Choice 开源权重 4.7B(活跃 4.5B) curateddiopenrouters1mbsom

Together AI 的决策模型。注意两个坑:只支持 Choice(返回单个选项字母),而且权重许可证至今未定稿(官方原文:release license being finalized)。

S1MB Task Avg
47.1
第 18 名 · 覆盖 137/137(100%)

模型信息

厂商
Together AI
类别
开源权重
参数规模
4.7B(活跃 4.5B)
基座模型
Qwen/Qwen3.5-4B
权重
可下载(开源)
许可证
许可证未定稿 需自行核对
输入模态
文本
决策原语
Choice
上下文
32,768 token
输入价格
$0.042/M tok
可微调
可以
微调方法
官方公开完整配方:Qwen3.5-4B 上 LoRA SFT(rank 8、1 epoch、lr 5e-5、2048 token),37,840 train / 4,568 val
微调硬件
官方称约 $17 可训一个自己的
训练方式
full fine-tune
OpenRouter 模型 ID
togethercomputer/tev1-4b-experimental
OpenRouter 计费
输入 $0.042/M tok · 输出 免费
上架状态
Early access

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
open jev / browser control v1 Choice 100.0 100.0 67.4 来源 ↗
open jev / customer control v1 Noul 100.0 100.0 58.0 来源 ↗
open jev / drone control v1 Noul 100.0 100.0 50.2 来源 ↗
open jev / mailroom control v1 Choice 100.0 100.0 82.9 来源 ↗
s1mb generalization contextual choice Choice 100.0 100.0 82.6 来源 ↗
dbpedia Choice 97.8 98.9 83.5 来源 ↗
bbq Choice 96.7 98.3 53.8 来源 ↗
open jev / citation control v1 Choice 95.5 100.0 40.9 来源 ↗
open jev / vizdoom basic v1 Noul 94.1 100.0 39.0 来源 ↗
arc Choice 92.7 98.5 64.4 来源 ↗
open jev / sponsor segment control v1 Choice 92.3 100.0 52.2 来源 ↗
snli Choice 90.3 98.4 46.0 来源 ↗
open jev / mailroom control v1 Noul 90.3 100.0 62.5 来源 ↗
open jev / phone extraction control v1 Noul 90.0 100.0 50.4 来源 ↗
s1mb generalization diverse choice Choice 88.1 100.0 69.5 来源 ↗
s1mb generalization contextual noul Noul 88.0 98.0 62.0 来源 ↗
s1mb generalization diverse noul Noul 88.0 100.0 59.2 来源 ↗
laya / ag news Choice 86.6 94.0 75.4 来源 ↗
open jev / email selection control v1 Noul 84.6 100.0 40.6 来源 ↗
laya / enron spam Noul 84.0 100.0 54.6 来源 ↗
scitail Noul 84.0 98.0 53.0 来源 ↗
open jev / customer control v1 Choice 82.6 82.6 64.3 来源 ↗
sdoh nli Noul 80.0 92.0 60.8 来源 ↗
openbookqa Choice 79.2 97.2 53.7 来源 ↗
quartz Choice 78.6 88.1 49.1 来源 ↗
massive Choice 77.5 96.7 64.5 来源 ↗
followir robust04 Noul 77.3 90.9 46.9 来源 ↗
impli Noul 76.0 88.0 46.7 来源 ↗
open jev / ir control v1 Choice 76.0 88.0 44.8 来源 ↗
open jev / entity alignment control v1 Noul 75.2 100.0 35.9 来源 ↗
open jev / reasoning control v1 Choice 75.0 95.0 28.4 来源 ↗
s1mb generalization diverse score Score 74.7 95.3 42.8 来源 ↗
mtop Choice 74.3 86.9 58.6 来源 ↗
open jev / workflow controls v1 / security incidents Noul 74.0 100.0 37.8 来源 ↗
open jev / silent failure control v1 Noul 74.0 100.0 38.8 来源 ↗
clinc Choice 73.1 99.3 57.3 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 72.9 100.0 34.6 来源 ↗
miqa Choice 72.7 95.5 61.9 来源 ↗
gretel pii Choice 72.1 94.5 43.1 来源 ↗
qasper Noul 70.0 88.0 41.6 来源 ↗
open jev / amount extraction control v1 Choice 68.6 100.0 55.0 来源 ↗
banking77 Choice 68.4 93.7 54.6 来源 ↗
hans Noul 68.0 100.0 43.3 来源 ↗
hatecheck Noul 68.0 98.0 47.5 来源 ↗
scone Choice 68.0 92.0 34.2 来源 ↗
defeasible nli Choice 67.4 80.4 32.8 来源 ↗
hwu64 Choice 66.0 92.5 56.8 来源 ↗
s1mb generalization contextual score Score 65.6 92.7 42.1 来源 ↗
open jev / vizdoom basic v1 Score 64.8 100.0 29.6 来源 ↗
open jev / reasoning control v1 Score 64.4 90.7 27.7 来源 ↗
paws Noul 64.0 88.0 45.7 来源 ↗
lonli Choice 63.9 86.9 39.2 来源 ↗
few nerd Choice 63.1 86.4 43.9 来源 ↗
open jev / phone extraction control v1 Choice 63.0 100.0 32.4 来源 ↗
open jev / workflow controls v1 / customer service Noul 62.3 100.0 41.6 来源 ↗
arct Choice 61.7 89.4 39.4 来源 ↗
argument quality Choice 61.2 95.9 37.4 来源 ↗
canttalk Noul 60.0 86.0 43.9 来源 ↗
crows pairs Choice 58.3 70.8 30.4 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 58.2 67.2 39.5 来源 ↗
ethos Noul 58.0 76.0 43.7 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 57.7 62.5 37.3 来源 ↗
open jev / ir control v1 Noul 57.1 100.0 47.3 来源 ↗
synthetic relevance / nanocoir Noul 56.5 73.0 33.6 来源 ↗
scicite Choice 56.5 71.0 44.0 来源 ↗
open jev / amount extraction control v1 Noul 56.3 100.0 33.3 来源 ↗
contract nli Choice 55.6 76.0 31.8 来源 ↗
creak Noul 52.0 88.0 38.8 来源 ↗
ethics Choice 50.0 83.3 20.9 来源 ↗
go emotions Noul 47.4 61.3 38.3 来源 ↗
laya / typed decisions Score 46.5 63.5 26.5 来源 ↗
babi nli Noul 46.0 84.0 37.0 来源 ↗
winobias Choice 45.6 91.3 30.1 来源 ↗
open jev / entity alignment control v1 Score 43.8 100.0 8.1 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 43.6 52.1 29.0 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 43.0 52.6 17.8 来源 ↗
aegis2 Noul 42.9 62.3 32.8 来源 ↗
open jev / painting geometry v1 Choice 42.1 100.0 46.2 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 42.0 44.9 17.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 42.0 54.0 28.6 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 40.9 44.3 20.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 37.9 46.7 15.2 来源 ↗
synthetic relevance / nanobeir / nanonq Score 37.5 49.4 9.3 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 36.7 46.2 25.1 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 36.1 47.0 26.6 来源 ↗
ruletaker Noul 36.0 96.0 27.7 来源 ↗
patent similarity Score 35.4 48.3 17.7 来源 ↗
lexcomp Noul 34.0 52.0 22.5 来源 ↗
open jev / customer control v1 Score 33.5 99.3 36.7 来源 ↗
open jev / reasoning control v1 Noul 33.5 100.0 25.3 来源 ↗
temporal nli Choice 32.3 80.7 18.3 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 31.8 38.6 10.7 来源 ↗
laya / typed decisions Noul 31.8 47.8 24.4 来源 ↗
followir core17 Noul 31.6 46.6 21.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 30.6 63.6 13.9 来源 ↗
sgd Choice 30.5 96.6 24.4 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 30.1 100.0 27.8 来源 ↗
open jev / painting geometry v1 Noul 29.5 100.0 27.2 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 29.1 36.0 7.7 来源 ↗
civil comments Noul 28.7 36.7 21.4 来源 ↗
logical entailment Noul 28.0 72.0 12.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 27.5 70.8 28.4 来源 ↗
stepgame Choice 27.4 82.1 17.8 来源 ↗
wiqa Choice 27.0 58.7 19.3 来源 ↗
ethics Noul 26.8 68.6 23.9 来源 ↗
gsm8k Choice 25.7 78.6 23.5 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 24.8 41.7 17.3 来源 ↗
winowhy Noul 24.0 62.0 18.8 来源 ↗
synthetic relevance / nanocoir Score 23.4 44.3 4.3 来源 ↗
laya / typed decisions Choice 23.1 34.5 18.9 来源 ↗
ud ewt Choice 22.4 56.7 13.2 来源 ↗
open jev / context retention control v1 Noul 22.0 100.0 19.0 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 21.9 36.7 7.2 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 20.7 48.2 15.6 来源 ↗
followir news21 Noul 20.7 42.7 14.9 来源 ↗
cladder Noul 18.0 94.0 17.7 来源 ↗
esci Choice 17.8 46.6 13.0 来源 ↗
robust lr Choice 16.7 66.7 18.8 来源 ↗
aqua rat Choice 12.7 77.5 10.7 来源 ↗
open jev / painting geometry v1 Score 12.6 100.0 15.2 来源 ↗
boardgameqa Choice 12.5 62.5 15.6 来源 ↗
fol nli Choice 9.8 54.1 10.4 来源 ↗
hh rlhf Choice 7.7 23.1 3.1 来源 ↗
open jev / snake v1 Choice 6.2 83.5 8.1 来源 ↗
open jev / email selection control v1 Choice 6.1 100.0 23.4 来源 ↗
open jev / snake v1 Noul 5.3 100.0 9.0 来源 ↗
tracie Noul 4.0 34.0 8.7 来源 ↗
spartqa Choice 3.2 38.7 4.4 来源 ↗
argument quality Noul 0.6 23.9 5.4 来源 ↗
corr2cause Choice 0.0 91.7 3.8 来源 ↗
nlsat Noul 0.0 22.0 2.1 来源 ↗
open jev / drone control v1 Choice 0.0 100.0 4.8 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
open jev / ir control v1 Score 0.0 98.6 10.1 来源 ↗
open jev / tic tac toe v1 Choice 0.0 48.5 3.5 来源 ↗
plane Noul 0.0 100.0 11.1 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗

同族条目(3)

型号参数Decision Index S1MBVision榜单
togethercomputer/Tev1-4B-experimental 当前 4.7B(活跃 4.5B) — 47.1 — s1mb
Together Tev1 Qwen3.5-4B 4.7B 34.1 — — di
Together Tev1 Qwen3.5-0.8B 873M 12.0 — — di

来自 systemonemodels.org 的详细介绍

What Tev1 is

Tev1 is Together AI’s fine-tune of Qwen3.5-4B into a Jev-style classifier: a System One model that reads a state and a closed set of options and returns one answer, rather than generated prose. Together published it as together/Tev1-4B-experimental on its serverless API and, alongside it, the full data recipe and a tutorial for training your own version. Unlike Jev’s typed Choice, Score and Noul questions, Tev1 has one output shape: given a state, a question and 2 to 24 labeled options, it returns the letter of the option it picked. The repo’s example script, decide.py, sends a support ticket and four possible intents and gets back a JSON object with the chosen label and key. There is no separate Score or Noul response format documented.

How it was built

Together fine-tuned Qwen/Qwen3.5-4B with LoRA (rank 8, one epoch, learning rate 5e-5, a 2,048-token sequence limit) on 37,840 to 38,340 training examples, depending on whether you read the GitHub README or the tutorial article; the two give slightly different totals for the same dataset. The mix draws from MultiNLI, BoolQ, Banking77, AG News and SST-5, plus generated policy, routing and research-classification examples. Together says the run cost about $17 and took roughly 25 minutes on their fine-tuning service. The README reports 880 of 1,000 correct on its main development set and 300 of 300 on a policy-transfer set. It is explicit that these are reused development benchmarks the team checked its own recipe against, not a held-out final test, and that endpoint access depends on your own Together account to reproduce.

What it’s good at

Short classification calls with a small, fixed option list: routing a support ticket to an intent, a yes/no/unsure read of a policy question, or picking a sentiment bucket, the same shape as the training datasets. Output tokens are free, so cost scales with the size of the state you send in. The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, gives an independent number. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Tev1-4B-experimental scores 29.24 and Tev1-0.8B-experimental 12.85, against Jev’s 57.91.

What it’s not for

Tev1 does not generate text, and it has no documented Score or Noul primitive, so anything needing a calibrated probability or a numeric level rather than a picked letter falls outside what’s published. The GitHub README calls its own benchmark numbers reused development results rather than an independent test, so treat any accuracy comparison with Jev as unverified.

Access today

The fine-tuning repo is MIT licensed and the weights are on Hugging Face under togethercomputer/Tev1-4B-experimental. The repo’s training and deployment scripts use Together’s fine-tuning and dedicated-endpoint services. The hosted together/Tev1-4B-experimental model is live on Together’s serverless API now, billed per the pricing above, with no waitlist mentioned in the launch materials. It is also listed on OpenRouter as togethercomputer/tev1-4b-experimental since 30 September 2026, served by Together at the same price.

Specifications

Question types Choice Max Choice options24 Score levelsNot documented Questions per callNot documented Total context2,048 tokens State budgetNot documented Rate limitNot documented SDKs The repo's README says the model takes a state, a question, and 2 to 24 options and returns one answer letter. The saved training recipe used a 2,048-token sequence limit; the maximum number of questions per call and any separate state-token budget are not documented. There is no distinct Score or Noul output: every answer is a selected option letter. OpenRouter accepts a 32,768-token context for the model, well above the 2,048-token training sequence limit, and sets a maximum of 8 completion tokens. Inputs past what training covered fall outside the tested range.

Versions

together/Tev1-4B-experimental, 23 Sep 2026, LoRA fine-tune of Qwen/Qwen3.5-4B, rank 8, one epoch, learning rate 5e-5. The repo's saved dev-set results: 880 of 1,000 on the main decision set and 300 of 300 on a policy-transfer set. The README says these are reused development benchmarks, not untouched final tests. Release notes

Use cases

What people use Tev1 for, one page per pattern. Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket. Choice Score Noul

Examples built with Tev1

The most-starred and most-viewed entries in the directory. Browse all examples.

Tev1 (Together AI): Jev-like classifier on Qwen3.5 4B

Data recipe and scripts to fine-tune a Jev-style classifier on Qwen3.5 4B from 37,840 examples for about $17. Together also serves the result, Tev1-4B-experimental, at $0.042 per million input tokens, and a step-by-step tutorial is on X. Together AI (@togethercompute)Open Tev1 (Together AI): Jev-like classifier on Qwen3.5 4B on github.com 23 starsTraining cost $17Serving price $0.042/M in Tool

Decision Index: Jev against 50+ open decision models

A leaderboard that runs Jev and 54 open System One models and clones through the same 120,000-question suite on one RTX PRO 6000, scoring accuracy and calibration. Jev leads the 0.2 edition at 51.67 on a chance-corrected scale, with AutoJev-27B close behind at 50.94. apolinario (poli) (@multimodalart) 43k viewsOpen Decision Index: Jev against 50+ open decision models on huggingface.co Article VendorTogether AI TypeCommercial, hosted API StatusEarly access Announced23 Sep 2026 AccessOpened 23 Sep 2026 Input price$0.042 per 1M tokens Output priceFree LatencyNot published Context2,048 tokens Rate limitNot published API docsapi.together.ai