CLM-8B
用对比目标给「状态 vs 候选动作」打分,而不是生成文本。对外提供 TypeSafe 兼容 API,并发布了数据配方与 scaling law 研究。
S1MB Task Avg
7.2
第 87 名 · 覆盖 137/137(100%)
Decision Index Full
6.2
第 101 名 · public 7.4
校准 ECE
0.323
越低越好 · Brier 0.904
模型信息
- 厂商
- Contrastive-LM
- 类别
- 开源权重
- 参数规模
- 7.6B(活跃 7.0B)
- 基座模型
- Qwen/Qwen3-8B-Base
- 权重
- 可下载(开源)
- 许可证
- 未标注 许可不明
- 输入模态
- 文本
- 决策原语
- ChoiceNoulScore
- 可微调
- 未确认
- 训练方式
- full fine-tune
- 上架状态
- Generally available
数据来源与链接
systemonemodels.orgHugging Face 权重GitHubbenchmarkheaven.commultimodalart-jev-decision-inde…X 帖子contrastive-lm.notion.site官方文档厂商页Decision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、s1mb、som。
分数与链接均指向原始出处。
Decision Index 领域得分
知识与推理
38.6
语言理解
42.8
检索与分类
12.2
工具与自动化
37.8
艺术与人类品味
35.0
S1MB · 137 个基准
Decision Index · 42 个基准
| 基准 | 类型 | 该模型 | 全场最佳 | 全场均值 | 对比 | 来源 |
|---|---|---|---|---|---|---|
| PhishNChips | 混合(Choice | 65.0 | 99.9 | 62.1 | 来源 ↗ | |
| NLI4CT | 混合(Choice | 56.5 | 86.2 | 69.3 | 来源 ↗ | |
| ARC-Easy | 混合(Choice | 56.4 | 99.5 | 87.3 | 来源 ↗ | |
| Humicroedit | 混合(Choice | 55.6 | 75.1 | 57.2 | 来源 ↗ | |
| CLadder | 混合(Choice | 53.9 | 97.7 | 62.5 | 来源 ↗ | |
| HoVer | 混合(Choice | 52.0 | 89.4 | 63.8 | 来源 ↗ | |
| RAGTruth | 混合(Choice | 51.8 | 86.0 | 52.1 | 来源 ↗ | |
| WinoGrande | 混合(Choice | 50.8 | 97.5 | 69.8 | 来源 ↗ | |
| cfcolor | 混合(Choice | 50.8 | 70.2 | 57.9 | 来源 ↗ | |
| CRUXEval | 混合(Choice | 42.8 | 87.9 | 50.7 | 来源 ↗ | |
| ARC-Challenge | 混合(Choice | 41.5 | 98.2 | 79.4 | 来源 ↗ | |
| HellaSwag | 混合(Choice | 40.4 | 98.5 | 73.0 | 来源 ↗ | |
| ToolRet | 混合(Choice | 40.3 | 69.1 | 55.1 | 来源 ↗ | |
| RouterBench | 混合(Choice | 39.3 | 80.1 | 73.2 | 来源 ↗ | |
| BFCL | 混合(Choice | 38.9 | 98.8 | 80.5 | 来源 ↗ | |
| When2Call | 混合(Choice | 38.7 | 91.7 | 57.8 | 来源 ↗ | |
| Habermas | 混合(Choice | 37.1 | 71.8 | 42.2 | 来源 ↗ | |
| MMLU | 混合(Choice | 35.8 | 91.9 | 64.4 | 来源 ↗ | |
| MuSR | 混合(Choice | 34.6 | 86.2 | 56.0 | 来源 ↗ | |
| FinEntity | 混合(Choice | 33.4 | 97.1 | 73.8 | 来源 ↗ | |
| BPoMP | 混合(Choice | 32.9 | 97.0 | 73.8 | 来源 ↗ | |
| ANLI | 混合(Choice | 31.6 | 98.6 | 54.0 | 来源 ↗ | |
| iSarcasmEval | 混合(Choice | 31.2 | 71.0 | 38.4 | 来源 ↗ | |
| BBH | 混合(Choice | 30.8 | 92.9 | 58.3 | 来源 ↗ | |
| BRIGHT | 混合(Choice | 29.1 | 50.9 | 36.1 | 来源 ↗ | |
| New Yorker | 混合(Choice | 25.4 | 82.0 | 52.9 | 来源 ↗ | |
| VAST | 混合(Choice | 22.7 | 82.0 | 50.1 | 来源 ↗ | |
| GPQA Diamond | 混合(Choice | 20.4 | 78.3 | 37.7 | 来源 ↗ | |
| ContractNLI | 混合(Choice | 20.1 | 86.5 | 59.9 | 来源 ↗ | |
| GSM8K | 混合(Choice | 17.5 | 83.5 | 36.0 | 来源 ↗ | |
| HLE | 混合(Choice | 15.8 | 20.4 | 12.3 | 来源 ↗ | |
| MMLU-Pro | 混合(Choice | 14.7 | 82.7 | 42.3 | 来源 ↗ | |
| SGD | 混合(Choice | 10.0 | 73.4 | 47.4 | 来源 ↗ | |
| CLINC150 | 混合(Choice | 8.9 | 97.4 | 67.8 | 来源 ↗ | |
| ChessBench | 混合(Choice | 8.2 | 40.2 | 14.2 | 来源 ↗ | |
| BANKING77 | 混合(Choice | 6.5 | 94.1 | 68.7 | 来源 ↗ | |
| API-Bank | 混合(Choice | 6.1 | 93.1 | 56.3 | 来源 ↗ | |
| SATA-Bench | 混合(Choice | 5.6 | 37.2 | 20.3 | 来源 ↗ | |
| Amazon ESCI | 混合(Choice | 5.0 | 61.6 | 40.9 | 来源 ↗ | |
| ACOS | 混合(Choice | 3.1 | 57.3 | 14.6 | 来源 ↗ | |
| Home appliances | 混合(Choice | 0.0 | 98.9 | 23.2 | 来源 ↗ | |
| POP909 | 混合(Choice | 0.0 | 74.6 | 12.6 | 来源 ↗ |
来自 systemonemodels.org 的详细介绍
What CLM is
CLM, short for Contrastive Language Model, is an open-weights System One model from a research group led by Jacky Kwok, whose X bio lists a Stanford CS PhD and Berkeley EECS background, with six co-authors credited in the repo’s citation. It reads a state and a set of candidate actions and scores each, rather than generating text. clm-serve exposes this over a POST /v1/systemone endpoint shaped like TypeSafe’s API, so it accepts Choice, Score and Noul questions the same way Jev does, plus a lower-level /v1/rank endpoint that scores any list of free-form candidates against a state.
How it is built
CLM trains two encoders, one for states and one for actions, with a contrastive objective (InfoNCE) that pulls a state’s embedding toward the action actually taken and away from the others. At inference, a typed question becomes a state plus a set of candidate action texts, and a softmax over the similarity scores is the answer distribution. The served model pairs a frozen Qwen3-8B backbone with a roughly 75MB trainable projection head, so most of the cost is one embedding per fresh text.
The authors describe three training stages: pre-training on about 60 million Nemotron question-answer pairs, mid-training on about 30 million synthetic hard negatives, and post-training on about 1 million agentic trajectories from public agent-trace datasets. The repo also documents scaling-law fits relating test contrastive loss to compute, model size and data size, published in full on the linked Notion blog.
What it’s good at, according to the authors
The authors report CLM-8B performs on par with Jev across computer-use, gaming and tool-calling tasks while running up to 9 times faster, with the largest gains when many candidate actions can be embedded once and reused. After lightweight fine-tuning, they report state-of-the-art verifier results on two agentic coding benchmarks: 87.6% on Terminal-Bench 2.1 and 81.6% on DeepSWE, both evaluated on held-out task sets, where they say Jev does not serve as an effective verifier. These are the authors’ own claims, run on their own benchmark harness, and have not been independently reproduced.
The disaggregated state and action encoders mean a fixed action set’s embeddings can be cached and reused as the state keeps changing, which the README frames as the main latency win for agent loops that revisit the same options.
The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, does not support the parity claim on its suite. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, CLM-v0.1-8B scores 7.40 against Jev’s 57.91. The index covers general decision tasks, not only the agent loops CLM targets.
What it’s not for
CLM’s own state and action encoders make it a two-embedding classifier, not a generative model, and its typed-question support depends on clm-serve translating Choice, Score and Noul requests into that shape rather than TypeSafe having defined the wire format. The repo requires a Linux host with an NVIDIA GPU to serve the reference encoder, so there is no documented CPU or Apple Silicon path.
Access today
Code and the CLM-8B weights are Apache 2.0 and hosted on GitHub and Hugging Face, with no waitlist. Running it means standing up a vllm serve pooling endpoint for the Qwen3-8B encoder plus clm-serve for the API and playground.
Specifications
Question types Choice Score Noul
Max Choice optionsNot documented
Score levelsNot documented
Questions per callNot documented
Total context2,048 tokens
State budgetNot documented
Rate limitNot documented
EndpointPOST /v1/systemone on your own server
SDKsPython: clm
The README's quickstart serves the Qwen3-8B encoder with --max-model-len 2048; the repo does not otherwise document a fixed cap on Choice options, Score levels, or questions per call. Score criteria must be an ordered list of at least 2 levels, matching TypeSafe's Score primitive. State can be a string, an object rendered as key: value text, or an array rendered as list lines, never JSON.
Versions
Contrastive-LM/CLM-v0.1-8B, 23 Sep 2026, The reference projection head served as clm-latest, about 75MB, trained against a Qwen3-8B backbone with last-token pooling. Apache 2.0 on Hugging Face. Release notes
Use cases
What people use CLM for, one page per pattern.
Workflow controlStarter
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Choice Score Noul
Real-time and agentsIntermediate
Agent routing and skill selection with System One models
An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess.
Choice Noul
Examples built with CLM
The most-starred and most-viewed entries in the directory. Browse all examples.
CLM: Contrastive Language Models
Open System One model that scores states against actions with a contrastive objective, serving CLM-8B behind a TypeSafe-compatible API. The authors report Jev-level results on computer-use, gaming and tool-calling with up to 9x lower latency, plus scaling laws and a fine-tuning guide.
Jacky Kwok (@jackyk02)Open CLM: Contrastive Language Models on github.com
215 starsvs Jev latency up to 9x fasterDeepSWE (tuned) 81.6%
Tool