Clef
27B 多模态决策模型,带视觉编码器,一轮前向并行给所有选项打分。是目前最强的「既有开源权重、又有官方微调服务」的决策模型。
S1MB Task Avg
56.0
第 7 名 · 覆盖 137/137(100%)
Decision Index Full
53.1
第 19 名 · public 61.7
校准 ECE
0.040
越低越好 · Brier 0.306
模型信息
- 厂商
- Cloudflare
- 类别
- 开源权重
- 参数规模
- 27B(活跃 26B)
- 基座模型
- Qwen/Qwen3.8-27B
- 权重
- 可下载(开源)
- 许可证
- Apache-2.0 可商用
- 输入模态
- 文本 图像 视频
- 决策原语
- ChoiceNoulScore
- 上下文
- 65,536 token
- 输入价格
- $0.24/M tok
- 延迟
- 209.3 ms 中位(Decision Index 实测)
- 可微调
- 可以
- 微调方法
- 冻结 backbone + rank-256 LoRA + joint schema head;label-smoothed CE + Brier loss 做校准,再加 RLCD。官方另提供 RL 微调服务(需申请 design partner)
- 微调硬件
- 单卡 H200 可推理;微调目前走官方 RL 服务
- 训练方式
- full fine-tune
- OpenRouter 模型 ID
cloudflare/clef- OpenRouter 计费
- 输入 $0.24/M tok · 输出 免费
- 上架状态
- Generally available
数据来源与链接
Hugging Face 权重官方博客官方文档OpenRoutersystemonemodels.orgclef-evals.workers-ai-mle.worke…cloudflare.comX 帖子厂商页Decision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、openrouter、s1mb、som。
分数与链接均指向原始出处。
Decision Index 领域得分
知识与推理
73.5
语言理解
86.5
检索与分类
74.3
工具与自动化
96.5
艺术与人类品味
63.2
S1MB · 137 个基准
Decision Index · 42 个基准
| 基准 | 类型 | 该模型 | 全场最佳 | 全场均值 | 对比 | 来源 |
|---|---|---|---|---|---|---|
| ARC-Easy | 混合(Choice | 99.0 | 99.5 | 87.3 | 来源 ↗ | |
| BFCL | 混合(Choice | 98.5 | 98.8 | 80.5 | 来源 ↗ | |
| HellaSwag | 混合(Choice | 98.2 | 98.5 | 73.0 | 来源 ↗ | |
| ARC-Challenge | 混合(Choice | 97.8 | 98.2 | 79.4 | 来源 ↗ | |
| CLINC150 | 混合(Choice | 97.4 | 97.4 | 67.8 | 来源 ↗ | |
| BPoMP | 混合(Choice | 97.0 | 97.0 | 73.8 | 来源 ↗ | |
| FinEntity | 混合(Choice | 96.1 | 97.1 | 73.8 | 来源 ↗ | |
| BANKING77 | 混合(Choice | 94.1 | 94.1 | 68.7 | 来源 ↗ | |
| CLadder | 混合(Choice | 94.1 | 97.7 | 62.5 | 来源 ↗ | |
| WinoGrande | 混合(Choice | 93.5 | 97.5 | 69.8 | 来源 ↗ | |
| API-Bank | 混合(Choice | 91.9 | 93.1 | 56.3 | 来源 ↗ | |
| MMLU | 混合(Choice | 90.2 | 91.9 | 64.4 | 来源 ↗ | |
| CRUXEval | 混合(Choice | 86.5 | 87.9 | 50.7 | 来源 ↗ | |
| MuSR | 混合(Choice | 83.8 | 86.2 | 56.0 | 来源 ↗ | |
| NLI4CT | 混合(Choice | 83.0 | 86.2 | 69.3 | 来源 ↗ | |
| ContractNLI | 混合(Choice | 81.3 | 86.5 | 59.9 | 来源 ↗ | |
| Home appliances | 混合(Choice | 80.7 | 98.9 | 23.2 | 来源 ↗ | |
| RouterBench | 混合(Choice | 79.7 | 80.1 | 73.2 | 来源 ↗ | |
| RAGTruth | 混合(Choice | 79.5 | 86.0 | 52.1 | 来源 ↗ | |
| PhishNChips | 混合(Choice | 79.3 | 99.9 | 62.1 | 来源 ↗ | |
| BBH | 混合(Choice | 73.6 | 92.9 | 58.3 | 来源 ↗ | |
| When2Call | 混合(Choice | 72.5 | 91.7 | 57.8 | 来源 ↗ | |
| New Yorker | 混合(Choice | 70.3 | 82.0 | 52.9 | 来源 ↗ | |
| ANLI | 混合(Choice | 70.0 | 98.6 | 54.0 | 来源 ↗ | |
| ToolRet | 混合(Choice | 69.1 | 69.1 | 55.1 | 来源 ↗ | |
| Habermas | 混合(Choice | 68.7 | 71.8 | 42.2 | 来源 ↗ | |
| Humicroedit | 混合(Choice | 66.5 | 75.1 | 57.2 | 来源 ↗ | |
| MMLU-Pro | 混合(Choice | 66.0 | 82.7 | 42.3 | 来源 ↗ | |
| cfcolor | 混合(Choice | 65.9 | 70.2 | 57.9 | 来源 ↗ | |
| HoVer | 混合(Choice | 65.6 | 89.4 | 63.8 | 来源 ↗ | |
| GSM8K | 混合(Choice | 59.9 | 83.5 | 36.0 | 来源 ↗ | |
| VAST | 混合(Choice | 59.3 | 82.0 | 50.1 | 来源 ↗ | |
| Amazon ESCI | 混合(Choice | 57.4 | 61.6 | 40.9 | 来源 ↗ | |
| GPQA Diamond | 混合(Choice | 48.5 | 78.3 | 37.7 | 来源 ↗ | |
| BRIGHT | 混合(Choice | 47.5 | 50.9 | 36.1 | 来源 ↗ | |
| iSarcasmEval | 混合(Choice | 46.4 | 71.0 | 38.4 | 来源 ↗ | |
| SGD | 混合(Choice | 43.8 | 73.4 | 47.4 | 来源 ↗ | |
| SATA-Bench | 混合(Choice | 33.5 | 37.2 | 20.3 | 来源 ↗ | |
| ACOS | 混合(Choice | 33.2 | 57.3 | 14.6 | 来源 ↗ | |
| ChessBench | 混合(Choice | 24.5 | 40.2 | 14.2 | 来源 ↗ | |
| POP909 | 混合(Choice | 15.9 | 74.6 | 12.6 | 来源 ↗ | |
| HLE | 混合(Choice | 12.8 | 20.4 | 12.3 | 来源 ↗ |
来自 systemonemodels.org 的详细介绍
What Clef is
Clef is a System One model trained by Cloudflare. You send a state and a set of typed questions, and it returns a probability for every allowed option of every question, with no generated text. It comes in two sizes. Clef is 27B parameters and Clef-flash is 9B. Both run on Cloudflare Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash.
Cloudflare calls the request format “fully Jev-API compatible”. It uses the same state and questions fields and the same Noul, Choice and Score question types as Jev, so the request body moves between the two with the model field changed.
What it adds to the Jev request
Two differences are Cloudflare’s own list. Clef reads images, where Jev reads text only. The docs add an optional images array of up to four images, which they call a Clef extension to the System One API. Its context window is 65,536 tokens, against 32,000 for Jev per Cloudflare’s post. The docs also say the state can be text, JSON, images or video.
How it was built
Cloudflare says each model keeps a frozen Qwen backbone, Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash. A routing head and rank-256 adapters are trained on top. Inference runs one forward pass over the input, then scores every valid answer in parallel, so no text is generated token by token. Training used label-smoothed cross-entropy with a Brier loss for calibration, synthetic data, and a second stage Cloudflare calls Reinforcement Learning for Calibrated Decisions. It builds on Cloudflare’s earlier DiffusionGemma experiment, which drew on Matt Mastracci’s djev. See how Jev is built for the comparison.
What Cloudflare reports
Every figure here is vendor-run: Cloudflare ran the evals and published the tables. On 10 tasks it picked from the Decision Index, Clef scores higher than Jev on eight and lower on two: When2Call, 72.37 against 80.97, and BRIGHT, 45.91 against 47.52. Clef-flash scores 66.77 on CLINC150+OOS against Jev’s 89.27. On TypeSafe’s workflow evals Clef beats Jev on invoice processing (64.7 against 61.8), customer service (76.3 against 76.0) and security incidents (62.9 against 61.7), and trails it on agent trace observability (68.5 against 71.6). Cloudflare also says Clef is currently the leader when evaluated against the Jev Decision Index. We have not seen an outside run.
What it’s good at
Short typed calls where you want a probability to act on: routing a support request, picking a team, scoring severity, or sending a call to a human when confidence is low. Cloudflare’s example classifies a website’s category from a fetched page. Because it reads images, it can also classify a screenshot or photo.
What it’s not for
It writes no text. Clef-flash scores well below Clef on some tasks, such as CLINC150+OOS (66.77 against 97.43, Cloudflare’s numbers). The latency figures are Cloudflare’s own and do not say how they were measured.
Access today
Both models are live on Workers AI, billed per input token. The weights are on Hugging Face under Apache 2.0, and Cloudflare also announced a fine-tuning service, hands-on at first and self-serve later. Cloudflare says it does not read, store or train on requests or responses unless you opt into fine-tuning.
Specifications
Question types Choice Score Noul
Max Choice optionsNot documented
Score levelsNot documented
Questions per call64
Total context65,536 tokens
State budgetNot documented
Rate limitNot published on the Workers AI model pages as of 2026-10-02.
EndpointPOST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef, or env.AI.run("@cf/cloudflare/clef", {...}) from a Worker
SDKs
Workers AI lists a 65,536-token context window for both models and says long text state is truncated to fit the token limit. A request takes 1 to 64 questions. The optional images array holds up to 4 PNG, JPEG or WebP images, each up to 4 MiB and 16 megapixels, 8 MiB in total, inside a 13 MiB request body. Remote image URLs are not accepted. The docs publish no cap on options per Choice or levels per Score.
Versions
@cf/cloudflare/clef, 1 Oct 2026, The 27B model, built on a frozen Qwen3.8-27B backbone. About 27.4B parameters per Hugging Face. Apache 2.0 weights as Cloudflare/clef. Release notes
@cf/cloudflare/clef-flash, 1 Oct 2026, The 9B model for latency-sensitive calls, built on a frozen Qwen3.5-9B backbone. About 9.4B parameters per Hugging Face. Apache 2.0 weights as Cloudflare/clef-flash. Release notes
Use cases
What people use Clef for, one page per pattern.
Workflow controlStarter
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Choice Score Noul
Workflow controlIntermediate
Intent and model routing with System One models
One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue.
Choice Score
Workflow controlIntermediate