decision.host
首页 / 模型 / CLM-8B

CLM-8B

Contrastive-LM ChoiceNoulScore 开源权重 7.6B(活跃 7.0B) curateddis1mbsom

用对比目标给「状态 vs 候选动作」打分,而不是生成文本。对外提供 TypeSafe 兼容 API,并发布了数据配方与 scaling law 研究。

S1MB Task Avg
7.2
第 87 名 · 覆盖 137/137(100%)
Decision Index Full
6.2
第 101 名 · public 7.4
校准 ECE
0.323
越低越好 · Brier 0.904

模型信息

厂商
Contrastive-LM
类别
开源权重
参数规模
7.6B(活跃 7.0B)
基座模型
Qwen/Qwen3-8B-Base
权重
可下载(开源)
许可证
未标注 许可不明
输入模态
文本
决策原语
ChoiceNoulScore
可微调
未确认
训练方式
full fine-tune
上架状态
Generally available

Decision Index 领域得分

知识与推理
38.6
语言理解
42.8
检索与分类
12.2
工具与自动化
37.8
艺术与人类品味
35.0

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
open jev / drone control v1 Noul 88.2 100.0 50.2 来源 ↗
open jev / ir control v1 Noul 78.6 100.0 47.3 来源 ↗
scitail Noul 70.0 98.0 53.0 来源 ↗
open jev / email selection control v1 Choice 53.0 100.0 23.4 来源 ↗
creak Noul 50.0 88.0 38.8 来源 ↗
open jev / mailroom control v1 Noul 48.3 100.0 62.5 来源 ↗
laya / enron spam Noul 48.0 100.0 54.6 来源 ↗
qasper Noul 44.0 88.0 41.6 来源 ↗
go emotions Noul 40.7 61.3 38.3 来源 ↗
open jev / reasoning control v1 Noul 40.4 100.0 25.3 来源 ↗
sdoh nli Noul 40.0 92.0 60.8 来源 ↗
open jev / drone control v1 Choice 38.6 100.0 4.8 来源 ↗
hatecheck Noul 34.0 98.0 47.5 来源 ↗
babi nli Noul 32.0 84.0 37.0 来源 ↗
open jev / phone extraction control v1 Noul 30.0 100.0 50.4 来源 ↗
ethos Noul 26.0 76.0 43.7 来源 ↗
open jev / amount extraction control v1 Choice 25.5 100.0 55.0 来源 ↗
open jev / mailroom control v1 Choice 25.0 100.0 82.9 来源 ↗
laya / ag news Choice 23.9 94.0 75.4 来源 ↗
s1mb generalization contextual choice Choice 22.8 100.0 82.6 来源 ↗
wiqa Choice 22.2 58.7 19.3 来源 ↗
followir robust04 Noul 21.9 90.9 46.9 来源 ↗
s1mb generalization diverse choice Choice 20.7 100.0 69.5 来源 ↗
civil comments Noul 19.5 36.7 21.4 来源 ↗
cladder Noul 18.0 94.0 17.7 来源 ↗
gsm8k Choice 17.1 78.6 23.5 来源 ↗
lexcomp Noul 16.0 52.0 22.5 来源 ↗
open jev / painting geometry v1 Score 15.1 100.0 15.2 来源 ↗
s1mb generalization contextual noul Noul 14.0 98.0 62.0 来源 ↗
massive Choice 13.9 96.7 64.5 来源 ↗
aegis2 Noul 13.3 62.3 32.8 来源 ↗
arc Choice 13.2 98.5 64.4 来源 ↗
open jev / customer control v1 Noul 12.5 100.0 58.0 来源 ↗
arct Choice 10.6 89.4 39.4 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 10.1 62.5 37.3 来源 ↗
bbq Choice 10.0 98.3 53.8 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 9.3 67.2 39.5 来源 ↗
esci Choice 8.2 46.6 13.0 来源 ↗
canttalk Noul 8.0 86.0 43.9 来源 ↗
s1mb generalization diverse score Score 7.7 95.3 42.8 来源 ↗
laya / typed decisions Noul 7.5 47.8 24.4 来源 ↗
dbpedia Choice 6.7 98.9 83.5 来源 ↗
ethics Noul 5.4 68.6 23.9 来源 ↗
mtop Choice 5.2 86.9 58.6 来源 ↗
miqa Choice 4.5 95.5 61.9 来源 ↗
followir core17 Noul 4.1 46.6 21.2 来源 ↗
laya / typed decisions Choice 4.1 34.5 18.9 来源 ↗
open jev / ir control v1 Choice 4.0 88.0 44.8 来源 ↗
s1mb generalization diverse noul Noul 4.0 100.0 59.2 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 3.9 46.2 25.1 来源 ↗
open jev / tic tac toe v1 Choice 3.0 48.5 3.5 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 2.8 41.7 17.3 来源 ↗
followir news21 Noul 2.7 42.7 14.9 来源 ↗
open jev / reasoning control v1 Score 2.3 90.7 27.7 来源 ↗
synthetic relevance / nanocoir Noul 2.2 73.0 33.6 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 2.2 70.8 28.4 来源 ↗
open jev / context retention control v1 Noul 2.0 100.0 19.0 来源 ↗
openbookqa Choice 1.4 97.2 53.7 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 0.5 52.1 29.0 来源 ↗
aqua rat Choice 0.0 77.5 10.7 来源 ↗
argument quality Choice 0.0 95.9 37.4 来源 ↗
argument quality Noul 0.0 23.9 5.4 来源 ↗
banking77 Choice 0.0 93.7 54.6 来源 ↗
boardgameqa Choice 0.0 62.5 15.6 来源 ↗
clinc Choice 0.0 99.3 57.3 来源 ↗
contract nli Choice 0.0 76.0 31.8 来源 ↗
corr2cause Choice 0.0 91.7 3.8 来源 ↗
crows pairs Choice 0.0 70.8 30.4 来源 ↗
defeasible nli Choice 0.0 80.4 32.8 来源 ↗
ethics Choice 0.0 83.3 20.9 来源 ↗
few nerd Choice 0.0 86.4 43.9 来源 ↗
fol nli Choice 0.0 54.1 10.4 来源 ↗
gretel pii Choice 0.0 94.5 43.1 来源 ↗
hans Noul 0.0 100.0 43.3 来源 ↗
hh rlhf Choice 0.0 23.1 3.1 来源 ↗
hwu64 Choice 0.0 92.5 56.8 来源 ↗
impli Noul 0.0 88.0 46.7 来源 ↗
laya / typed decisions Score 0.0 63.5 26.5 来源 ↗
logical entailment Noul 0.0 72.0 12.2 来源 ↗
lonli Choice 0.0 86.9 39.2 来源 ↗
nlsat Noul 0.0 22.0 2.1 来源 ↗
open jev / amount extraction control v1 Noul 0.0 100.0 33.3 来源 ↗
open jev / browser control v1 Choice 0.0 100.0 67.4 来源 ↗
open jev / citation control v1 Choice 0.0 100.0 40.9 来源 ↗
open jev / customer control v1 Choice 0.0 82.6 64.3 来源 ↗
open jev / customer control v1 Score 0.0 99.3 36.7 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
open jev / email selection control v1 Noul 0.0 100.0 40.6 来源 ↗
open jev / entity alignment control v1 Noul 0.0 100.0 35.9 来源 ↗
open jev / entity alignment control v1 Score 0.0 100.0 8.1 来源 ↗
open jev / ir control v1 Score 0.0 98.6 10.1 来源 ↗
open jev / painting geometry v1 Choice 0.0 100.0 46.2 来源 ↗
open jev / painting geometry v1 Noul 0.0 100.0 27.2 来源 ↗
open jev / phone extraction control v1 Choice 0.0 100.0 32.4 来源 ↗
open jev / reasoning control v1 Choice 0.0 95.0 28.4 来源 ↗
open jev / silent failure control v1 Noul 0.0 100.0 38.8 来源 ↗
open jev / snake v1 Choice 0.0 83.5 8.1 来源 ↗
open jev / snake v1 Noul 0.0 100.0 9.0 来源 ↗
open jev / sponsor segment control v1 Choice 0.0 100.0 52.2 来源 ↗
open jev / vizdoom basic v1 Noul 0.0 100.0 39.0 来源 ↗
open jev / vizdoom basic v1 Score 0.0 100.0 29.6 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 0.0 100.0 27.8 来源 ↗
open jev / workflow controls v1 / customer service Noul 0.0 100.0 41.6 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 0.0 100.0 34.6 来源 ↗
open jev / workflow controls v1 / security incidents Noul 0.0 100.0 37.8 来源 ↗
patent similarity Score 0.0 48.3 17.7 来源 ↗
paws Noul 0.0 88.0 45.7 来源 ↗
plane Noul 0.0 100.0 11.1 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗
quartz Choice 0.0 88.1 49.1 来源 ↗
robust lr Choice 0.0 66.7 18.8 来源 ↗
ruletaker Noul 0.0 96.0 27.7 来源 ↗
s1mb generalization contextual score Score 0.0 92.7 42.1 来源 ↗
scicite Choice 0.0 71.0 44.0 来源 ↗
scone Choice 0.0 92.0 34.2 来源 ↗
sgd Choice 0.0 96.6 24.4 来源 ↗
snli Choice 0.0 98.4 46.0 来源 ↗
spartqa Choice 0.0 38.7 4.4 来源 ↗
stepgame Choice 0.0 82.1 17.8 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 0.0 48.2 15.6 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 0.0 36.0 7.7 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 0.0 44.9 17.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 0.0 54.0 28.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 0.0 46.7 15.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 0.0 63.6 13.9 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 0.0 52.6 17.8 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 0.0 38.6 10.7 来源 ↗
synthetic relevance / nanobeir / nanonq Score 0.0 49.4 9.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 0.0 36.7 7.2 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 0.0 47.0 26.6 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 0.0 44.3 20.6 来源 ↗
synthetic relevance / nanocoir Score 0.0 44.3 4.3 来源 ↗
temporal nli Choice 0.0 80.7 18.3 来源 ↗
tracie Noul 0.0 34.0 8.7 来源 ↗
ud ewt Choice 0.0 56.7 13.2 来源 ↗
winobias Choice 0.0 91.3 30.1 来源 ↗
winowhy Noul 0.0 62.0 18.8 来源 ↗

Decision Index · 42 个基准

基准类型该模型全场最佳 全场均值对比来源
PhishNChips 混合(Choice 65.0 99.9 62.1 来源 ↗
NLI4CT 混合(Choice 56.5 86.2 69.3 来源 ↗
ARC-Easy 混合(Choice 56.4 99.5 87.3 来源 ↗
Humicroedit 混合(Choice 55.6 75.1 57.2 来源 ↗
CLadder 混合(Choice 53.9 97.7 62.5 来源 ↗
HoVer 混合(Choice 52.0 89.4 63.8 来源 ↗
RAGTruth 混合(Choice 51.8 86.0 52.1 来源 ↗
WinoGrande 混合(Choice 50.8 97.5 69.8 来源 ↗
cfcolor 混合(Choice 50.8 70.2 57.9 来源 ↗
CRUXEval 混合(Choice 42.8 87.9 50.7 来源 ↗
ARC-Challenge 混合(Choice 41.5 98.2 79.4 来源 ↗
HellaSwag 混合(Choice 40.4 98.5 73.0 来源 ↗
ToolRet 混合(Choice 40.3 69.1 55.1 来源 ↗
RouterBench 混合(Choice 39.3 80.1 73.2 来源 ↗
BFCL 混合(Choice 38.9 98.8 80.5 来源 ↗
When2Call 混合(Choice 38.7 91.7 57.8 来源 ↗
Habermas 混合(Choice 37.1 71.8 42.2 来源 ↗
MMLU 混合(Choice 35.8 91.9 64.4 来源 ↗
MuSR 混合(Choice 34.6 86.2 56.0 来源 ↗
FinEntity 混合(Choice 33.4 97.1 73.8 来源 ↗
BPoMP 混合(Choice 32.9 97.0 73.8 来源 ↗
ANLI 混合(Choice 31.6 98.6 54.0 来源 ↗
iSarcasmEval 混合(Choice 31.2 71.0 38.4 来源 ↗
BBH 混合(Choice 30.8 92.9 58.3 来源 ↗
BRIGHT 混合(Choice 29.1 50.9 36.1 来源 ↗
New Yorker 混合(Choice 25.4 82.0 52.9 来源 ↗
VAST 混合(Choice 22.7 82.0 50.1 来源 ↗
GPQA Diamond 混合(Choice 20.4 78.3 37.7 来源 ↗
ContractNLI 混合(Choice 20.1 86.5 59.9 来源 ↗
GSM8K 混合(Choice 17.5 83.5 36.0 来源 ↗
HLE 混合(Choice 15.8 20.4 12.3 来源 ↗
MMLU-Pro 混合(Choice 14.7 82.7 42.3 来源 ↗
SGD 混合(Choice 10.0 73.4 47.4 来源 ↗
CLINC150 混合(Choice 8.9 97.4 67.8 来源 ↗
ChessBench 混合(Choice 8.2 40.2 14.2 来源 ↗
BANKING77 混合(Choice 6.5 94.1 68.7 来源 ↗
API-Bank 混合(Choice 6.1 93.1 56.3 来源 ↗
SATA-Bench 混合(Choice 5.6 37.2 20.3 来源 ↗
Amazon ESCI 混合(Choice 5.0 61.6 40.9 来源 ↗
ACOS 混合(Choice 3.1 57.3 14.6 来源 ↗
Home appliances 混合(Choice 0.0 98.9 23.2 来源 ↗
POP909 混合(Choice 0.0 74.6 12.6 来源 ↗

来自 systemonemodels.org 的详细介绍

What CLM is

CLM, short for Contrastive Language Model, is an open-weights System One model from a research group led by Jacky Kwok, whose X bio lists a Stanford CS PhD and Berkeley EECS background, with six co-authors credited in the repo’s citation. It reads a state and a set of candidate actions and scores each, rather than generating text. clm-serve exposes this over a POST /v1/systemone endpoint shaped like TypeSafe’s API, so it accepts Choice, Score and Noul questions the same way Jev does, plus a lower-level /v1/rank endpoint that scores any list of free-form candidates against a state.

How it is built

CLM trains two encoders, one for states and one for actions, with a contrastive objective (InfoNCE) that pulls a state’s embedding toward the action actually taken and away from the others. At inference, a typed question becomes a state plus a set of candidate action texts, and a softmax over the similarity scores is the answer distribution. The served model pairs a frozen Qwen3-8B backbone with a roughly 75MB trainable projection head, so most of the cost is one embedding per fresh text. The authors describe three training stages: pre-training on about 60 million Nemotron question-answer pairs, mid-training on about 30 million synthetic hard negatives, and post-training on about 1 million agentic trajectories from public agent-trace datasets. The repo also documents scaling-law fits relating test contrastive loss to compute, model size and data size, published in full on the linked Notion blog.

What it’s good at, according to the authors

The authors report CLM-8B performs on par with Jev across computer-use, gaming and tool-calling tasks while running up to 9 times faster, with the largest gains when many candidate actions can be embedded once and reused. After lightweight fine-tuning, they report state-of-the-art verifier results on two agentic coding benchmarks: 87.6% on Terminal-Bench 2.1 and 81.6% on DeepSWE, both evaluated on held-out task sets, where they say Jev does not serve as an effective verifier. These are the authors’ own claims, run on their own benchmark harness, and have not been independently reproduced. The disaggregated state and action encoders mean a fixed action set’s embeddings can be cached and reused as the state keeps changing, which the README frames as the main latency win for agent loops that revisit the same options. The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, does not support the parity claim on its suite. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, CLM-v0.1-8B scores 7.40 against Jev’s 57.91. The index covers general decision tasks, not only the agent loops CLM targets.

What it’s not for

CLM’s own state and action encoders make it a two-embedding classifier, not a generative model, and its typed-question support depends on clm-serve translating Choice, Score and Noul requests into that shape rather than TypeSafe having defined the wire format. The repo requires a Linux host with an NVIDIA GPU to serve the reference encoder, so there is no documented CPU or Apple Silicon path.

Access today

Code and the CLM-8B weights are Apache 2.0 and hosted on GitHub and Hugging Face, with no waitlist. Running it means standing up a vllm serve pooling endpoint for the Qwen3-8B encoder plus clm-serve for the API and playground.

Specifications

Question types Choice Score Noul Max Choice optionsNot documented Score levelsNot documented Questions per callNot documented Total context2,048 tokens State budgetNot documented Rate limitNot documented EndpointPOST /v1/systemone on your own server SDKsPython: clm The README's quickstart serves the Qwen3-8B encoder with --max-model-len 2048; the repo does not otherwise document a fixed cap on Choice options, Score levels, or questions per call. Score criteria must be an ordered list of at least 2 levels, matching TypeSafe's Score primitive. State can be a string, an object rendered as key: value text, or an array rendered as list lines, never JSON.

Versions

Contrastive-LM/CLM-v0.1-8B, 23 Sep 2026, The reference projection head served as clm-latest, about 75MB, trained against a Qwen3-8B backbone with last-token pooling. Apache 2.0 on Hugging Face. Release notes

Use cases

What people use CLM for, one page per pattern. Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket. Choice Score Noul Real-time and agentsIntermediate

Agent routing and skill selection with System One models

An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess. Choice Noul

Examples built with CLM

The most-starred and most-viewed entries in the directory. Browse all examples.

CLM: Contrastive Language Models

Open System One model that scores states against actions with a contrastive objective, serving CLM-8B behind a TypeSafe-compatible API. The authors report Jev-level results on computer-use, gaming and tool-calling with up to 9x lower latency, plus scaling laws and a fine-tuning guide. Jacky Kwok (@jackyk02)Open CLM: Contrastive Language Models on github.com 215 starsvs Jev latency up to 9x fasterDeepSWE (tuned) 81.6% Tool