kev-0.5b
覆盖 0.5B / 0.6B / 0.8B / 4B / 8B / 9B / 27B 的模型家族,Apache-2.0,TypeSafe 兼容 API,4B 同时在 OpenRouter 上以 $0.042/百万提供。如果你要微调自己的决策模型,这是入门首选。
S1MB Task Avg
10.1
第 78 名 · 覆盖 137/137(100%)
模型信息
- 厂商
- Jared Palmer
- 类别
- 开源权重
- 参数规模
- 4.2B(活跃 358M)
- 基座模型
- Qwen3.5 / Qwen3.8 系列
- 权重
- 可下载(开源)
- 许可证
- Apache-2.0 可商用
- 输入模态
- 文本
- 决策原语
- ChoiceNoulScore
- 上下文
- 8,192 token
- 输入价格
- $0.042/M tok
- 延迟
- 41.5-145.2 ms
- 可微调
- 可以
- 微调方法
- 文档最完整的开源决策模型:koev.train --init_from <ckpt>;LoRA + pointer head;逐 checkpoint 拟合温度(0.8B T=2.35 / 9B T=2.30)
- 微调硬件
- 0.8B bf16 仅需 4 GB 显存;4B 走 Modal 约 $1/H100
- 训练方式
- full fine-tune
- OpenRouter 模型 ID
jaredpalmer/kev-4b- OpenRouter 计费
- 输入 $0.042/M tok · 输出 免费
- 上架状态
- Generally available
数据来源与链接
代码仓库HF · 4BHF · 9BHF · 27BOpenRoutersystemonemodels.orgHugging Face 权重pypi.org官方博客clef-evals.workers-ai-mle.worke…GitHubX 帖子benchmarkheaven.commultimodalart-jev-decision-inde…官方文档厂商页Decision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、openrouter、s1mb、som。
分数与链接均指向原始出处。
S1MB · 137 个基准
同族条目(9)
| 型号 | 参数 | Decision Index | S1MB | Vision | 榜单 |
|---|---|---|---|---|---|
| kev-0.5b 当前 | 4.2B(活跃 358M) | — | 10.1 | — | s1mb |
| Kev 27B | 28B | 58.8 | — | — | di |
| jaredpalmer/kev-9b | 7.9B(活跃 6.9B) | 43.3 | 42.4 | — | s1mb, di |
| Kev 4B r10 | 4.7B | 39.5 | — | — | di |
| Kev 0.8B r15 | 870M | 13.7 | — | — | di |
| kev-0.6b | 597M(活跃 441M) | — | 14.1 | — | s1mb |
| jaredpalmer/kev-0.8b | 753M(活跃 499M) | — | 18.7 | — | s1mb |
| jaredpalmer/kev-4b | 4.2B(活跃 3.6B) | — | 42.3 | — | s1mb |
| kev-8b | 7.6B(活跃 6.9B) | — | 36.0 | — | s1mb |
来自 systemonemodels.org 的详细介绍
What Kev is
Kev is a family of open-source System One models published by Jared Palmer. It reads a state and answers typed Choice, Score and Noul questions about it with probabilities, the same three shapes Jev uses. You download the weights from Hugging Face and run them on CUDA, ROCm or Apple Silicon through MLX, or call Kev-4B on OpenRouter. The 27B model needs an 80 GB data-centre GPU and has no Mac path.
The server exposes POST /v1/systemone with the same request and response shape as TypeSafe’s API, so the typesafe-sdk Python client works once you point base_url at your own machine. The repo also ships a Modal script that deploys Kev-4B to an L40S behind a bearer key, and a Hugging Face Space runs Kev-4B and Kev-0.8B in the browser.
OpenRouter added Kev-4B on 25 September 2026, hosted by SiliconFlow in fp8. It costs $0.042 per million input tokens with free output and an 8,192-token context, the same price as Jev. OpenRouter says it takes the same request as Jev, so switching means changing the model field to jaredpalmer/kev-4b. The 0.8B, 9B and 27B models are not hosted there.
How it is built
Each checkpoint is a rank-16 LoRA adapter plus a small pointer head on a Qwen base: Qwen3.5 at 0.8B, 4B and 9B, and Qwen3.8 at 27B. The head scores each option against a decision token at the end of its question, and a softmax turns those scores into probabilities. Questions share the state but cannot see each other.
Training is plain cross-entropy on 10,000 examples from ten public datasets, plus 896 generated policy examples and 1,680 from generated rule structures. The README says no Jev outputs were used. Kev-27B trains for one epoch on Kev-9B’s data, plus 1,400 records where the question sits among 1k to 6k tokens of unrelated text, and it uses soft targets where the right answer is genuinely ambiguous. Each checkpoint ships with a temperature fitted on in-distribution data, which changes the probabilities but never the answer.
The 0.8B, 4B and 9B models got a short second training pass on generated examples on 2026-09-21, with the earlier weights kept at revision v7-base. Kev-4B was updated again on 2026-09-24 after one epoch on consumer-finance complaints.
Kev-27B differs from the others in its base. The smaller models start from Qwen base checkpoints. Kev-27B starts from Qwen3.8-27B, Qwen’s post-trained release, and the README says its training data is unknown. Comparisons with Jev or with the smaller Kevs therefore do not test the method alone.
What its own evals show
These are the author’s numbers from the repo’s own benchmark, which also runs the Jev comparison. On datasets Kev was not trained on, Kev-9B scores 0.822 on the development set against Jev’s 0.857, and 0.852 on a test set Jev has not been run on. Kev-27B scores 0.848 on the development set, within a point of Jev, and 0.896 on the test set. On held-out examples from its training sources Kev-9B scores 0.872 against Jev’s 0.845, and Kev-27B 0.866. The README says this is not a controlled comparison, because Jev’s training data is unknown, and Kev-27B’s post-trained base adds a second unknown.
The first independent test is the Decision Index 0.2.1 by multimodalart, updated 28 September 2026. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Jev scores 57.91, Kev-9B 38.48, Kev-4B 34.64 and Kev-0.8B 14.60. Kev-9B ranks 26th of the 70 entries, 19.4 points behind Jev. Their expected calibration error, the average gap between stated confidence and actual accuracy, is 0.138 for Kev-9B, 0.176 for Kev-4B and 0.074 for Kev-0.8B, against Jev’s 0.074. Kev-27B was not in the index on 30 September.
Speed is self-reported too. Kev-4B on an L40S takes 41.5ms to 145ms of model time for a new state, depending on how many questions you ask and how long the state is. Kev-27B takes 46.5ms to 178ms on a B200 and 75ms to 278ms on an H100. The README’s sample request on an Apple M5 in bf16 came back in 495ms.
What it is not for
Knowledge questions are the widest gap: Kev-9B scores 0.74 on MMLU and Kev-27B 0.84, where Jev scores 0.90. Changing the order of options can change an answer. Training covered at most 384 state tokens, so long documents fall outside it even though the server accepts a 65,536-token state. The server handles one request at a time, and the Qwen3.5 models run slowly on a Mac. Kev-27B needs an 80 GB GPU, with 55 GB of weights and about 66 GB in use once the serving buffers are counted, and it has no Mac path.
Specifications
Question types Choice Score Noul
Max Choice options255
Score levelsUp to 255
Questions per callNot documented
Total context8,192 tokens
State budgetNot documented
Rate limitNone when self-hosted, where the server handles one request at a time. OpenRouter does not publish a rate limit for Kev-4B.
EndpointPOST /v1/systemone on your own server, or Kev-4B through OpenRouter
SDKsPython: typesafe-sdk
The README says the server accepts a state of up to 65,536 tokens plus 8,192 for each question, but training used at most 384 state tokens and 1,024 for the state plus one question, so longer inputs fall outside what training covered. Kev-27B holds up better on long text: 0.833 on questions buried in 1k to 6k tokens of unrelated text, against 0.556 for Kev-9B. OpenRouter lists an 8,192-token context for Kev-4B. Choice takes 1 to 255 options and Score 1 to 255 levels. A request can carry any number of questions; the server runs them a 16,384-token row at a time, so memory does not grow with the question count.
Versions
jaredpalmer/kev-0.8b, 21 Sep 2026, LoRA adapter and pointer head on Qwen3.5-0.8B-Base. Weights updated on 2026-09-21 with a second training pass on generated examples; the previous weights are at revision v7-base. Release notes
jaredpalmer/kev-4b, 24 Sep 2026, The recommended starting point, on Qwen3.5-4B-Base. Updated on 2026-09-21 like the others, then again on 2026-09-24 with one epoch on 5,219 consumer-finance complaint narratives; the previous weights are at revision night2-du-release. OpenRouter serves this build as kev-4b-20260924. Release notes
jaredpalmer/kev-9b, 21 Sep 2026, The most accurate Kev on a Qwen3.5 base. Updated on 2026-09-21; the previous weights are at revision v7-base. Release notes
jaredpalmer/kev-27b, 24 Sep 2026, The most accurate Kev, a LoRA adapter and pointer head on Qwen3.8-27B, Qwen's post-trained release rather than a base checkpoint. Apache 2.0. It needs an 80 GB GPU (55 GB of weights in bf16, about 66 GB with serving buffers) and has no Mac path. Release notes
Use cases
What people use Kev for, one page per pattern.
Workflow controlStarter
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Choice Score Noul
Workflow controlIntermediate
Confidence-gated actions with System One models
Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars.
Choice Score
Examples built with Kev
The most-starred and most-viewed entries in the directory. Browse all examples.
Introducing Clef: our open-source decision models, and new RL fine-tuning platform
Cloudflare's launch post for Clef and Clef-flash, two decision models on Workers AI that accept the Jev request format and open weights under Apache 2.0. It covers the architecture, Cloudflare's own benchmark tables against Jev, Kev and Laya, and a new reinforcement learning service.
Michelle Chen (@michellechen) 140k viewsOpen Introducing Clef: our open-source decision models, and new RL fine-tuning platform on blog.cloudflare.com
Article Support inbox triage with System One models
Decision Model Leaderboard
Cloudflare's live leaderboard for decision models on the Decision Index suite, with Clef, Clef-flash, Jev, Kev-9B and Laya plotted by score and latency. Cloudflare built and ran it, so treat the ranking as vendor-run.
Cloudflare (@ritakozlov) 41k viewsOpen Decision Model Leaderboard on clef-evals.workers-ai-mle.workers.dev
Article