decision.host
首页 / 模型 / Clef

Clef

Cloudflare ChoiceNoulScore图像视频 开源权重 27B(活跃 26B) curateddiopenrouters1mbsom

27B 多模态决策模型,带视觉编码器,一轮前向并行给所有选项打分。是目前最强的「既有开源权重、又有官方微调服务」的决策模型。

S1MB Task Avg
56.0
第 7 名 · 覆盖 137/137(100%)
Decision Index Full
53.1
第 19 名 · public 61.7
校准 ECE
0.040
越低越好 · Brier 0.306

模型信息

厂商
Cloudflare
类别
开源权重
参数规模
27B(活跃 26B)
基座模型
Qwen/Qwen3.8-27B
权重
可下载(开源)
许可证
Apache-2.0 可商用
输入模态
文本 图像 视频
决策原语
ChoiceNoulScore
上下文
65,536 token
输入价格
$0.24/M tok
延迟
209.3 ms 中位(Decision Index 实测)
可微调
可以
微调方法
冻结 backbone + rank-256 LoRA + joint schema head;label-smoothed CE + Brier loss 做校准,再加 RLCD。官方另提供 RL 微调服务(需申请 design partner)
微调硬件
单卡 H200 可推理;微调目前走官方 RL 服务
训练方式
full fine-tune
OpenRouter 模型 ID
cloudflare/clef
OpenRouter 计费
输入 $0.24/M tok · 输出 免费
上架状态
Generally available

数据来源与链接

本站聚合自: curated、di、openrouter、s1mb、som。 分数与链接均指向原始出处。

Decision Index 领域得分

知识与推理
73.5
语言理解
86.5
检索与分类
74.3
工具与自动化
96.5
艺术与人类品味
63.2

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
open jev / browser control v1 Choice 100.0 100.0 67.4 来源 ↗
open jev / customer control v1 Noul 100.0 100.0 58.0 来源 ↗
open jev / email selection control v1 Choice 100.0 100.0 23.4 来源 ↗
open jev / email selection control v1 Noul 100.0 100.0 40.6 来源 ↗
open jev / entity alignment control v1 Noul 100.0 100.0 35.9 来源 ↗
open jev / mailroom control v1 Choice 100.0 100.0 82.9 来源 ↗
open jev / painting geometry v1 Choice 100.0 100.0 46.2 来源 ↗
open jev / phone extraction control v1 Noul 100.0 100.0 50.4 来源 ↗
dbpedia Choice 98.9 98.9 83.5 来源 ↗
s1mb generalization diverse choice Choice 98.7 100.0 69.5 来源 ↗
clinc Choice 98.6 99.3 57.3 来源 ↗
open jev / sponsor segment control v1 Choice 97.4 100.0 52.2 来源 ↗
s1mb generalization contextual choice Choice 97.4 100.0 82.6 来源 ↗
arc Choice 97.1 98.5 64.4 来源 ↗
s1mb generalization diverse noul Noul 96.0 100.0 59.2 来源 ↗
massive Choice 94.5 96.7 64.5 来源 ↗
open jev / vizdoom basic v1 Noul 94.1 100.0 39.0 来源 ↗
open jev / citation control v1 Choice 93.9 100.0 40.9 来源 ↗
banking77 Choice 93.7 93.7 54.6 来源 ↗
miqa Choice 93.2 95.5 61.9 来源 ↗
openbookqa Choice 93.1 97.2 53.7 来源 ↗
laya / enron spam Noul 92.0 100.0 54.6 来源 ↗
open jev / context retention control v1 Noul 92.0 100.0 19.0 来源 ↗
hwu64 Choice 91.5 92.5 56.8 来源 ↗
open jev / mailroom control v1 Noul 90.6 100.0 62.5 来源 ↗
cladder Noul 90.0 94.0 17.7 来源 ↗
s1mb generalization contextual noul Noul 90.0 98.0 62.0 来源 ↗
argument quality Choice 89.8 95.9 37.4 来源 ↗
laya / ag news Choice 88.1 94.0 75.4 来源 ↗
hans Noul 88.0 100.0 43.3 来源 ↗
open jev / silent failure control v1 Noul 88.0 100.0 38.8 来源 ↗
sdoh nli Noul 88.0 92.0 60.8 来源 ↗
sgd Choice 86.4 96.6 24.4 来源 ↗
open jev / ir control v1 Noul 85.7 100.0 47.3 来源 ↗
mtop Choice 84.0 86.9 58.6 来源 ↗
open jev / ir control v1 Choice 84.0 88.0 44.8 来源 ↗
scitail Noul 84.0 98.0 53.0 来源 ↗
open jev / painting geometry v1 Score 82.7 100.0 15.2 来源 ↗
s1mb generalization diverse score Score 82.7 95.3 42.8 来源 ↗
winobias Choice 82.6 91.3 30.1 来源 ↗
open jev / drone control v1 Noul 82.3 100.0 50.2 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 82.0 100.0 34.6 来源 ↗
bbq Choice 81.7 98.3 53.8 来源 ↗
open jev / workflow controls v1 / security incidents Noul 81.0 100.0 37.8 来源 ↗
open jev / amount extraction control v1 Choice 80.4 100.0 55.0 来源 ↗
open jev / customer control v1 Choice 78.3 82.6 64.3 来源 ↗
few nerd Choice 77.1 86.4 43.9 来源 ↗
lonli Choice 77.0 86.9 39.2 来源 ↗
quartz Choice 76.2 88.1 49.1 来源 ↗
qasper Noul 76.0 88.0 41.6 来源 ↗
scone Choice 76.0 92.0 34.2 来源 ↗
gretel pii Choice 76.0 94.5 43.1 来源 ↗
contract nli Choice 75.9 76.0 31.8 来源 ↗
snli Choice 75.8 98.4 46.0 来源 ↗
gsm8k Choice 75.7 78.6 23.5 来源 ↗
s1mb generalization contextual score Score 74.5 92.7 42.1 来源 ↗
arct Choice 72.3 89.4 39.4 来源 ↗
creak Noul 72.0 88.0 38.8 来源 ↗
open jev / vizdoom basic v1 Score 72.0 100.0 29.6 来源 ↗
open jev / amount extraction control v1 Noul 71.9 100.0 33.3 来源 ↗
babi nli Noul 70.0 84.0 37.0 来源 ↗
open jev / reasoning control v1 Choice 70.0 95.0 28.4 来源 ↗
followir robust04 Noul 69.8 90.9 46.9 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 68.7 70.8 28.4 来源 ↗
ethics Noul 68.1 68.6 23.9 来源 ↗
scicite Choice 67.7 71.0 44.0 来源 ↗
defeasible nli Choice 67.4 80.4 32.8 来源 ↗
ethics Choice 66.7 83.3 20.9 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 66.4 100.0 27.8 来源 ↗
hatecheck Noul 66.0 98.0 47.5 来源 ↗
paws Noul 66.0 88.0 45.7 来源 ↗
open jev / phone extraction control v1 Choice 65.2 100.0 32.4 来源 ↗
impli Noul 64.0 88.0 46.7 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 63.1 67.2 39.5 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 62.5 62.5 37.3 来源 ↗
open jev / reasoning control v1 Score 60.8 90.7 27.7 来源 ↗
open jev / workflow controls v1 / customer service Noul 60.6 100.0 41.6 来源 ↗
crows pairs Choice 60.4 70.8 30.4 来源 ↗
ruletaker Noul 60.0 96.0 27.7 来源 ↗
robust lr Choice 58.3 66.7 18.8 来源 ↗
synthetic relevance / nanocoir Noul 58.0 73.0 33.6 来源 ↗
open jev / painting geometry v1 Noul 57.1 100.0 27.2 来源 ↗
laya / typed decisions Score 56.2 63.5 26.5 来源 ↗
canttalk Noul 56.0 86.0 43.9 来源 ↗
open jev / customer control v1 Score 55.7 99.3 36.7 来源 ↗
go emotions Noul 54.8 61.3 38.3 来源 ↗
open jev / ir control v1 Score 54.4 98.6 10.1 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 54.0 54.0 28.6 来源 ↗
aegis2 Noul 53.6 62.3 32.8 来源 ↗
aqua rat Choice 49.3 77.5 10.7 来源 ↗
winowhy Noul 48.0 62.0 18.8 来源 ↗
ud ewt Choice 46.7 56.7 13.2 来源 ↗
boardgameqa Choice 45.3 62.5 15.6 来源 ↗
esci Choice 45.2 46.6 13.0 来源 ↗
ethos Noul 44.0 76.0 43.7 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 44.0 47.0 26.6 来源 ↗
laya / typed decisions Noul 42.2 47.8 24.4 来源 ↗
lexcomp Noul 42.0 52.0 22.5 来源 ↗
temporal nli Choice 40.3 80.7 18.3 来源 ↗
open jev / entity alignment control v1 Score 39.7 100.0 8.1 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 39.3 44.3 20.6 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 39.2 52.1 29.0 来源 ↗
patent similarity Score 38.8 48.3 17.7 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 38.2 52.6 17.8 来源 ↗
followir core17 Noul 37.1 46.6 21.2 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 36.2 48.2 15.6 来源 ↗
fol nli Choice 32.8 54.1 10.4 来源 ↗
civil comments Noul 31.8 36.7 21.4 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 30.9 44.9 17.6 来源 ↗
open jev / snake v1 Noul 30.4 100.0 9.0 来源 ↗
wiqa Choice 30.2 58.7 19.3 来源 ↗
laya / typed decisions Choice 29.8 34.5 18.9 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 29.8 46.7 15.2 来源 ↗
open jev / reasoning control v1 Noul 29.0 100.0 25.3 来源 ↗
stepgame Choice 27.4 82.1 17.8 来源 ↗
followir news21 Noul 26.7 42.7 14.9 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 25.3 63.6 13.9 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 25.0 46.2 25.1 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 24.1 41.7 17.3 来源 ↗
argument quality Noul 23.9 23.9 5.4 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 23.4 38.6 10.7 来源 ↗
tracie Noul 18.0 34.0 8.7 来源 ↗
spartqa Choice 16.1 38.7 4.4 来源 ↗
logical entailment Noul 14.0 72.0 12.2 来源 ↗
plane Noul 14.0 100.0 11.1 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 11.5 36.0 7.7 来源 ↗
open jev / snake v1 Choice 9.3 83.5 8.1 来源 ↗
corr2cause Choice 6.7 91.7 3.8 来源 ↗
open jev / tic tac toe v1 Choice 5.3 48.5 3.5 来源 ↗
hh rlhf Choice 5.1 23.1 3.1 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 4.5 36.7 7.2 来源 ↗
nlsat Noul 4.0 22.0 2.1 来源 ↗
synthetic relevance / nanobeir / nanonq Score 0.3 49.4 9.3 来源 ↗
open jev / drone control v1 Choice 0.0 100.0 4.8 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗
synthetic relevance / nanocoir Score 0.0 44.3 4.3 来源 ↗

Decision Index · 42 个基准

基准类型该模型全场最佳 全场均值对比来源
ARC-Easy 混合(Choice 99.0 99.5 87.3 来源 ↗
BFCL 混合(Choice 98.5 98.8 80.5 来源 ↗
HellaSwag 混合(Choice 98.2 98.5 73.0 来源 ↗
ARC-Challenge 混合(Choice 97.8 98.2 79.4 来源 ↗
CLINC150 混合(Choice 97.4 97.4 67.8 来源 ↗
BPoMP 混合(Choice 97.0 97.0 73.8 来源 ↗
FinEntity 混合(Choice 96.1 97.1 73.8 来源 ↗
BANKING77 混合(Choice 94.1 94.1 68.7 来源 ↗
CLadder 混合(Choice 94.1 97.7 62.5 来源 ↗
WinoGrande 混合(Choice 93.5 97.5 69.8 来源 ↗
API-Bank 混合(Choice 91.9 93.1 56.3 来源 ↗
MMLU 混合(Choice 90.2 91.9 64.4 来源 ↗
CRUXEval 混合(Choice 86.5 87.9 50.7 来源 ↗
MuSR 混合(Choice 83.8 86.2 56.0 来源 ↗
NLI4CT 混合(Choice 83.0 86.2 69.3 来源 ↗
ContractNLI 混合(Choice 81.3 86.5 59.9 来源 ↗
Home appliances 混合(Choice 80.7 98.9 23.2 来源 ↗
RouterBench 混合(Choice 79.7 80.1 73.2 来源 ↗
RAGTruth 混合(Choice 79.5 86.0 52.1 来源 ↗
PhishNChips 混合(Choice 79.3 99.9 62.1 来源 ↗
BBH 混合(Choice 73.6 92.9 58.3 来源 ↗
When2Call 混合(Choice 72.5 91.7 57.8 来源 ↗
New Yorker 混合(Choice 70.3 82.0 52.9 来源 ↗
ANLI 混合(Choice 70.0 98.6 54.0 来源 ↗
ToolRet 混合(Choice 69.1 69.1 55.1 来源 ↗
Habermas 混合(Choice 68.7 71.8 42.2 来源 ↗
Humicroedit 混合(Choice 66.5 75.1 57.2 来源 ↗
MMLU-Pro 混合(Choice 66.0 82.7 42.3 来源 ↗
cfcolor 混合(Choice 65.9 70.2 57.9 来源 ↗
HoVer 混合(Choice 65.6 89.4 63.8 来源 ↗
GSM8K 混合(Choice 59.9 83.5 36.0 来源 ↗
VAST 混合(Choice 59.3 82.0 50.1 来源 ↗
Amazon ESCI 混合(Choice 57.4 61.6 40.9 来源 ↗
GPQA Diamond 混合(Choice 48.5 78.3 37.7 来源 ↗
BRIGHT 混合(Choice 47.5 50.9 36.1 来源 ↗
iSarcasmEval 混合(Choice 46.4 71.0 38.4 来源 ↗
SGD 混合(Choice 43.8 73.4 47.4 来源 ↗
SATA-Bench 混合(Choice 33.5 37.2 20.3 来源 ↗
ACOS 混合(Choice 33.2 57.3 14.6 来源 ↗
ChessBench 混合(Choice 24.5 40.2 14.2 来源 ↗
POP909 混合(Choice 15.9 74.6 12.6 来源 ↗
HLE 混合(Choice 12.8 20.4 12.3 来源 ↗

来自 systemonemodels.org 的详细介绍

What Clef is

Clef is a System One model trained by Cloudflare. You send a state and a set of typed questions, and it returns a probability for every allowed option of every question, with no generated text. It comes in two sizes. Clef is 27B parameters and Clef-flash is 9B. Both run on Cloudflare Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Cloudflare calls the request format “fully Jev-API compatible”. It uses the same state and questions fields and the same Noul, Choice and Score question types as Jev, so the request body moves between the two with the model field changed.

What it adds to the Jev request

Two differences are Cloudflare’s own list. Clef reads images, where Jev reads text only. The docs add an optional images array of up to four images, which they call a Clef extension to the System One API. Its context window is 65,536 tokens, against 32,000 for Jev per Cloudflare’s post. The docs also say the state can be text, JSON, images or video.

How it was built

Cloudflare says each model keeps a frozen Qwen backbone, Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash. A routing head and rank-256 adapters are trained on top. Inference runs one forward pass over the input, then scores every valid answer in parallel, so no text is generated token by token. Training used label-smoothed cross-entropy with a Brier loss for calibration, synthetic data, and a second stage Cloudflare calls Reinforcement Learning for Calibrated Decisions. It builds on Cloudflare’s earlier DiffusionGemma experiment, which drew on Matt Mastracci’s djev. See how Jev is built for the comparison.

What Cloudflare reports

Every figure here is vendor-run: Cloudflare ran the evals and published the tables. On 10 tasks it picked from the Decision Index, Clef scores higher than Jev on eight and lower on two: When2Call, 72.37 against 80.97, and BRIGHT, 45.91 against 47.52. Clef-flash scores 66.77 on CLINC150+OOS against Jev’s 89.27. On TypeSafe’s workflow evals Clef beats Jev on invoice processing (64.7 against 61.8), customer service (76.3 against 76.0) and security incidents (62.9 against 61.7), and trails it on agent trace observability (68.5 against 71.6). Cloudflare also says Clef is currently the leader when evaluated against the Jev Decision Index. We have not seen an outside run.

What it’s good at

Short typed calls where you want a probability to act on: routing a support request, picking a team, scoring severity, or sending a call to a human when confidence is low. Cloudflare’s example classifies a website’s category from a fetched page. Because it reads images, it can also classify a screenshot or photo.

What it’s not for

It writes no text. Clef-flash scores well below Clef on some tasks, such as CLINC150+OOS (66.77 against 97.43, Cloudflare’s numbers). The latency figures are Cloudflare’s own and do not say how they were measured.

Access today

Both models are live on Workers AI, billed per input token. The weights are on Hugging Face under Apache 2.0, and Cloudflare also announced a fine-tuning service, hands-on at first and self-serve later. Cloudflare says it does not read, store or train on requests or responses unless you opt into fine-tuning.

Specifications

Question types Choice Score Noul Max Choice optionsNot documented Score levelsNot documented Questions per call64 Total context65,536 tokens State budgetNot documented Rate limitNot published on the Workers AI model pages as of 2026-10-02. EndpointPOST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef, or env.AI.run("@cf/cloudflare/clef", {...}) from a Worker SDKs Workers AI lists a 65,536-token context window for both models and says long text state is truncated to fit the token limit. A request takes 1 to 64 questions. The optional images array holds up to 4 PNG, JPEG or WebP images, each up to 4 MiB and 16 megapixels, 8 MiB in total, inside a 13 MiB request body. Remote image URLs are not accepted. The docs publish no cap on options per Choice or levels per Score.

Versions

@cf/cloudflare/clef, 1 Oct 2026, The 27B model, built on a frozen Qwen3.8-27B backbone. About 27.4B parameters per Hugging Face. Apache 2.0 weights as Cloudflare/clef. Release notes @cf/cloudflare/clef-flash, 1 Oct 2026, The 9B model for latency-sensitive calls, built on a frozen Qwen3.5-9B backbone. About 9.4B parameters per Hugging Face. Apache 2.0 weights as Cloudflare/clef-flash. Release notes

Use cases

What people use Clef for, one page per pattern. Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket. Choice Score Noul Workflow controlIntermediate

Intent and model routing with System One models

One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue. Choice Score Workflow controlIntermediate