togethercomputer/Tev1-4B-experimental
Together AI 的决策模型。注意两个坑:只支持 Choice(返回单个选项字母),而且权重许可证至今未定稿(官方原文:release license being finalized)。
S1MB Task Avg
47.1
第 18 名 · 覆盖 137/137(100%)
模型信息
- 厂商
- Together AI
- 类别
- 开源权重
- 参数规模
- 4.7B(活跃 4.5B)
- 基座模型
- Qwen/Qwen3.5-4B
- 权重
- 可下载(开源)
- 许可证
- 许可证未定稿 需自行核对
- 输入模态
- 文本
- 决策原语
- Choice
- 上下文
- 32,768 token
- 输入价格
- $0.042/M tok
- 可微调
- 可以
- 微调方法
- 官方公开完整配方:Qwen3.5-4B 上 LoRA SFT(rank 8、1 epoch、lr 5e-5、2048 token),37,840 train / 4,568 val
- 微调硬件
- 官方称约 $17 可训一个自己的
- 训练方式
- full fine-tune
- OpenRouter 模型 ID
togethercomputer/tev1-4b-experimental- OpenRouter 计费
- 输入 $0.042/M tok · 输出 免费
- 上架状态
- Early access
数据来源与链接
Hugging Face 权重代码仓库OpenRoutersystemonemodels.orgGitHubtogether.aiapi.together.aimultimodalart-jev-decision-inde…X 帖子官方文档厂商页Decision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、openrouter、s1mb、som。
分数与链接均指向原始出处。
S1MB · 137 个基准
同族条目(3)
| 型号 | 参数 | Decision Index | S1MB | Vision | 榜单 |
|---|---|---|---|---|---|
| togethercomputer/Tev1-4B-experimental 当前 | 4.7B(活跃 4.5B) | — | 47.1 | — | s1mb |
| Together Tev1 Qwen3.5-4B | 4.7B | 34.1 | — | — | di |
| Together Tev1 Qwen3.5-0.8B | 873M | 12.0 | — | — | di |
来自 systemonemodels.org 的详细介绍
What Tev1 is
Tev1 is Together AI’s fine-tune of Qwen3.5-4B into a Jev-style classifier: a System One model that reads a state and a closed set of options and returns one answer, rather than generated prose. Together published it as together/Tev1-4B-experimental on its serverless API and, alongside it, the full data recipe and a tutorial for training your own version.
Unlike Jev’s typed Choice, Score and Noul questions, Tev1 has one output shape: given a state, a question and 2 to 24 labeled options, it returns the letter of the option it picked. The repo’s example script, decide.py, sends a support ticket and four possible intents and gets back a JSON object with the chosen label and key. There is no separate Score or Noul response format documented.
How it was built
Together fine-tuned Qwen/Qwen3.5-4B with LoRA (rank 8, one epoch, learning rate 5e-5, a 2,048-token sequence limit) on 37,840 to 38,340 training examples, depending on whether you read the GitHub README or the tutorial article; the two give slightly different totals for the same dataset. The mix draws from MultiNLI, BoolQ, Banking77, AG News and SST-5, plus generated policy, routing and research-classification examples. Together says the run cost about $17 and took roughly 25 minutes on their fine-tuning service.
The README reports 880 of 1,000 correct on its main development set and 300 of 300 on a policy-transfer set. It is explicit that these are reused development benchmarks the team checked its own recipe against, not a held-out final test, and that endpoint access depends on your own Together account to reproduce.
What it’s good at
Short classification calls with a small, fixed option list: routing a support ticket to an intent, a yes/no/unsure read of a policy question, or picking a sentiment bucket, the same shape as the training datasets. Output tokens are free, so cost scales with the size of the state you send in.
The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, gives an independent number. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Tev1-4B-experimental scores 29.24 and Tev1-0.8B-experimental 12.85, against Jev’s 57.91.
What it’s not for
Tev1 does not generate text, and it has no documented Score or Noul primitive, so anything needing a calibrated probability or a numeric level rather than a picked letter falls outside what’s published. The GitHub README calls its own benchmark numbers reused development results rather than an independent test, so treat any accuracy comparison with Jev as unverified.
Access today
The fine-tuning repo is MIT licensed and the weights are on Hugging Face under togethercomputer/Tev1-4B-experimental. The repo’s training and deployment scripts use Together’s fine-tuning and dedicated-endpoint services. The hosted together/Tev1-4B-experimental model is live on Together’s serverless API now, billed per the pricing above, with no waitlist mentioned in the launch materials. It is also listed on OpenRouter as togethercomputer/tev1-4b-experimental since 30 September 2026, served by Together at the same price.
Specifications
Question types Choice
Max Choice options24
Score levelsNot documented
Questions per callNot documented
Total context2,048 tokens
State budgetNot documented
Rate limitNot documented
SDKs
The repo's README says the model takes a state, a question, and 2 to 24 options and returns one answer letter. The saved training recipe used a 2,048-token sequence limit; the maximum number of questions per call and any separate state-token budget are not documented. There is no distinct Score or Noul output: every answer is a selected option letter. OpenRouter accepts a 32,768-token context for the model, well above the 2,048-token training sequence limit, and sets a maximum of 8 completion tokens. Inputs past what training covered fall outside the tested range.
Versions
together/Tev1-4B-experimental, 23 Sep 2026, LoRA fine-tune of Qwen/Qwen3.5-4B, rank 8, one epoch, learning rate 5e-5. The repo's saved dev-set results: 880 of 1,000 on the main decision set and 300 of 300 on a policy-transfer set. The README says these are reused development benchmarks, not untouched final tests. Release notes
Use cases
What people use Tev1 for, one page per pattern.
Workflow controlStarter
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Choice Score Noul
Examples built with Tev1
The most-starred and most-viewed entries in the directory. Browse all examples.
Tev1 (Together AI): Jev-like classifier on Qwen3.5 4B
Data recipe and scripts to fine-tune a Jev-style classifier on Qwen3.5 4B from 37,840 examples for about $17. Together also serves the result, Tev1-4B-experimental, at $0.042 per million input tokens, and a step-by-step tutorial is on X.
Together AI (@togethercompute)Open Tev1 (Together AI): Jev-like classifier on Qwen3.5 4B on github.com
23 starsTraining cost $17Serving price $0.042/M in
Tool
Decision Index: Jev against 50+ open decision models
A leaderboard that runs Jev and 54 open System One models and clones through the same 120,000-question suite on one RTX PRO 6000, scoring accuracy and calibration. Jev leads the 0.2 edition at 51.67 on a chance-corrected scale, with AutoJev-27B close behind at 50.94.
apolinario (poli) (@multimodalart) 43k viewsOpen Decision Index: Jev against 50+ open decision models on huggingface.co
Article
VendorTogether AI
TypeCommercial, hosted API
StatusEarly access
Announced23 Sep 2026
AccessOpened 23 Sep 2026
Input price$0.042 per 1M tokens
Output priceFree
LatencyNot published
Context2,048 tokens
Rate limitNot published
API docsapi.together.ai