decision.host

决策模型(System One)介绍

一种只做判断、不写文章的模型:输入 state 与类型化问题,输出概率分布。 这一页解释它的三种原语、与 LLM 的区别、统一 API 的现状,以及全部术语。

什么是决策模型

决策模型(Decision Model,也叫 System One 模型或类型化概率决策模型)是一类只做判断、不写文章的模型。 你给它一段 state(文本、JSON,部分模型还支持图像)和若干类型化问题, 它返回概率分布:每个选项一个概率,外加一个 confidence。它不生成解释、不做推理链、不输出自由文本—— 判断之后的动作由你的代码决定。

这个品类的起点是 2026 年 9 月 15 日 TypeSafe AI 发布的 Jev 1.13。 它在两周内催生了 100 多个开源复刻、两套公开评测(S1MB 与 Decision Index), 以及 OpenAI、Perplexity、Cloudflare、Liquid 等厂商各自的对标产品。 决策模型要替代的是「问一个大模型一个窄问题、再从它的文本里抠出标签」这一步: 输出可被程序直接分支、概率可校准可设阈值、一次前向传播所以更快也更便宜。

三种原语

Choice 选择
从你定义的选项里选一个
返回:选中项 + 每个选项的概率 + 一个 confidence
限制:Jev 上最多 255 个选项
典型用途:意图分类、工单分派、工具选择、模型路由
例:这封工单该交给哪个团队?
Noul 是否
某个条件是否成立
返回:一个 0–1 的概率(P(是))
限制:Jev 上不返回独立 confidence
典型用途:风控、内容审核、验证、门控
例:客户是在描述一个软件缺陷吗?
Score 打分
按你写的有序标准打分
返回:概率加权分数 + 每档概率 + legend
限制:Jev 上 2–10 档
典型用途:严重度、质量评分、优先级、置信路由
例:这个问题有多严重?

本站定位

决策模型在 2026 年 9 月 15 日 Jev 发布之后进入高速爆发期:两周内出现 100 多个开源复刻、两套公开评测、多家大厂的对标产品,以及 arXiv 上每天新增的相关论文。信息增长的速度,已经超过任何人靠追更能跟上的程度。

这里做的事情是「汇总与统一」:把散落在 Hugging Face、arXiv、厂商文档、第三方目录里的模型、评测、论文和工程经验收进同一个数据库,用同一套口径呈现,并保留每一条数据的原始出处。目标不是做一个更好的榜单,而是让任何人在十分钟内建立对这类模型的完整认知,并在需要时直达一手资料。

「决策模型」或许叫直觉模型更贴切。

观点
它并不做决策——它做的是直觉判断。更贴切的名字是「直觉模型」。

决策包含权衡、后果与责任;而这类模型做的事情是:在你给定的选项上,凭已经内化的知识给出一个概率分布,快速指一个「最像对」的答案。它不推演、不比较代价、不承担后果。就过程而言,它更接近人「拍脑袋」的那一下,而不是「想清楚之后再做决定」。

按双过程理论的说法,它对应的是 System 1——快速、直觉、省力;而不是 System 2 那种慢速、可追溯的推理。TypeSafe 自己把这个家族命名为 System One,其实比「决策模型」这个中文译法更准确。

名字重要,因为它决定了人们期待什么。叫「决策模型」,会让人以为它可以替你把事决定下来;叫「直觉模型」,你就知道它给的是一个先验判断,该不该照它执行,仍然是你的责任。

它让本来就难解释的大模型,更难解释了

观点
拿掉了生成过程,也就拿掉了唯一能被人读的那部分。

大语言模型的可解释性本来就是一笔糊涂账:注意力权重、神经元激活、思维链文本,没有一个是可靠的因果解释。但至少思维链给了人一个读得下去的东西——哪怕它可能只是事后编出来的理由。

决策模型把这一层直接删掉了:不生成任何文本,只返回一个概率分布。结果是更难解释。概率可以校准(说 0.9 就真的九成对),但「为什么是 0.9」无处可问——没有文本、没有中间态、没有可回放的过程。

所以工程上的合理位置是:低风险、可回滚、有人兜底的判断点。用它的概率做分流与阈值门控,把不确定的那部分交给人或交给更强的模型,而不是把它当作不可质疑的结论。

这个用法早就有,代价也很早就暴露了

工程经验
要求模型只输出 ABCD,确实不到一秒;但没有了 think,质量差很多。

在 Jev 出现之前,很多团队在业务里已经用过同类做法:把要求和答案结构写进上下文,要求模型只输出一个选项字母。这样做的收益非常直接——不需要解析自然语言、不需要处理格式漂移、延迟压到亚秒级、成本几乎可以忽略。

但问题同样直接:没有了思考过程,回答质量明显更差。同一道判断,让模型先写出理由再给标签,往往比直接吐标签更准;把推理过程拿掉,等于主动放弃了模型在困难样本上的那部分能力。

这正是决策模型的取舍:用质量上界换速度与成本。它适合那些本来就不需要深想的判断(这条消息该归到哪个类别、这个截图里有没有异常、这段话是否违规);而不适合把一道需要推理的问题硬压成一次前向。选型时先问一句:这件事人来做需要想吗?需要想,就别交给直觉模型。

为什么是现在爆发

趋势判断
需求一直在,缺的是行业共识;接口一旦统一,剩下的就是工程量。

这类需求并不新鲜,过去十年里以各种形式存在过:文本分类、意图识别、内容审核、rerank、规则引擎加模型打分。真正缺的不是技术,而是行业共识——可做的事情太多、收益分散在无数个具体场景里,没有人愿意为一个标准投入。

Jev 做对的一件事,是把接口形态固定下来:state + 类型化问题(Choice / Noul / Score)→ 概率分布,输出免费、按输入计费。标准一旦清晰,生态就会自己长出来:两周内 100 多个开源复刻、两套公开评测(S1MB 与 Decision Index)、Cloudflare 开源 Clef 并配套 RL 微调、OpenAI 与 Perplexity 各自推出 Decisions 端点。

共识形成之后,剩下的就是工程与数据——所以接下来会进入快车道。本站的存在意义也在于此:这段路上每天都有新东西,需要有人把它收拢、对齐、留档。

各家实现差异很大,本质是同一件事

观点
都是利用大语言模型已经内化的知识,做一次快速的、可能不那么精确的判断。

从外面看,这些模型差异极大:有在冻结的 27B 主干上加一个 head 的,有在 400M 编码器上重头训练的;有读标签 token 的 logit 的,有把选项拼进输入做成对打分的;有单次前向的,也有走受限解码的。尺寸从 17M 到 35B,延迟从 5 毫秒到 500 毫秒。

但如果问「它凭什么能判断」,答案只有一个:底模在预训练阶段已经把这些判断所需的知识内化进去了。决策模型做的事情,是把这些知识在一次前向里读出来,而不是重新学一遍。这也解释了为什么:

换句话说,架构决定的是「读得多准」,数据决定的是「读得到什么」。具体流派见下方「实现方式探索」。

  • 小模型在熟悉领域能逼近大模型,换个领域就崩——它读出的是基座里已有的东西;
  • 合成数据比参数量更关键——数据决定的是「读出哪些知识」;
  • 它快——判断所需的计算在执行第一个 token 前就基本完成了。

多模态是更有价值的方向

趋势判断
生产里真正需要判断的东西,大多数带着图。

文本判断点(意图分类、情感、NLI)虽然数量庞大,但很多已经被传统小模型以更低的成本解决了。而带图的判断点——截图、单据、图表、监控画面、网页布局、商品图片——往往没有现成的小模型可用,过去只能靠人工、或者靠昂贵的通用大模型硬扛。

这正是多模态决策模型的机会。而且从数据上看,它已经在领先:Decision Index 的 Vision 子榜最高分 69.8,高于文本主榜的 62.8;已经有 4 个托管端点支持图像输入(Cloudflare Clef / Clef-flash、OpenAI gpt-6-luna-decisions、Perplexity pplx-decider-v1.1-27b),另有 Intern-Decision、decider-2b-vision、JPT 系列等开源选择。

本质是降本增效,终点是业务场景的后训练

观点
先用最强的模型把判断做对,再用它的输出把便宜模型教对。

决策模型的价值不在「更聪明」,而在同一件事更便宜、更快:输入价 $0.02–0.24/百万 token、输出 token 免费、一次前向。省下来的是推理算力、是延迟、是解析与重试的工程成本。

但通用决策模型只解决「起步」问题。真正落地时,你会发现业务判断的长尾极其琐碎:你们公司特有的分类体系、你们自己的合规口径、你们那个行业的术语。这些不会出现在任何公开训练数据里。

于是路径变得清晰:用高阶模型在你自己的数据上跑出结果,再用这些结果去训练或校准一个小决策模型。本质接近蒸馏,只不过蒸出来的不是文本生成能力,而是判断——这正是 arXiv 上那篇《LLM-as-Jev: LLMs Are Already Jev-Style Decision Models — When and How to Fine-Tune Them》在讨论的事情。

所以一个实用的判断标准是:开源且提供完整后训练路径的模型,长期价值更高。只能调 API 的闭源端点适合快速验证,而你要沉淀的是自己的判断能力——那部分资产必须能落在自己的权重上。

按这个标准,目前值得优先研究的开源权重是:Cloudflare Clef / Clef-flash(Apache-2.0,有官方 RL 微调服务)、Perplexity pplx-decider-v1-27b(Apache-2.0,约 49 GiB,可自行 LoRA),以及带完整训练配方的 kev(0.8B 只要 4 GB 显存)、von、laya、decider、bekko(单张 5090 可训)、minojev(只训 0.8M head,笔记本 37 分钟)。多模态场景可以照 jev-smol 的路线:SmolVLM2-500M + 语言塔 LoRA,在 Apple Silicon 上 35 分钟出第一版。

实现方式探索

同样叫「决策模型」,各家在「怎么把判断读出来」这件事上差异很大。下面按读出架构、训练路径、校准手法三层拆开。这些差异决定了模型的速度、校准质量,以及它能不能被你继续后训练。

五种读出架构

固定槽位头Fixed slots
怎么做:用一个固定宽度的输出层替换原来的语言模型头,每个槽位对应一类问题的答案(例如 256 路选项槽 + 三档分数槽)。
优点:完全不生成 token,延迟最低;输出空间在训练时就确定,不依赖 prompt 渲染。
代价:选项上限写死在权重里;换 schema 必须重训或至少重新校准。
代表实现:OpenThai-SystemOne(256 路槽位替换 LM head)
选项标记Option markers
怎么做:把每个选项的标签当作一段特殊文本拼在候选位置,取该位置上的表示或 logit 作为分数。可以配合顺序无关的注意力掩码。
优点:天然支持可变数量的选项;对选项文本的措辞敏感,可以靠写清楚 criteria 提升准确率。
代价:选项顺序、措辞与数量会影响结果,需要专门做顺序与措辞的鲁棒性测试。
代表实现:Laya(两层 option-marker 打分器,另有 act/escalate 头)、von(option-marker 头 + 顺序无关注意力掩码)、Kev(LoRA + pointer head 读选项 logit)、Decision-1.0(共享候选头,读候选端点与 query 向量)
标签 logit 读出Label logits
怎么做:不新增结构,直接把选项字母/标签 token 放在答案位置做 teacher forcing,取该 token 的 log-prob,再对同题所有选项做 softmax。
优点:实现最轻,只需一次前向;概率直接来自模型分布,配合温度拟合即可校准;任何指令微调过的 LLM 都能改造。
代价:受 tokenizer 影响(多 token 标签要取平均);需要严格保证训练与推理的 prompt 逐字节一致。
代表实现:Bespoke Nimble(直接对允许答案 token 打分,T=2.179)、decider(letter-logit 读出后除以存好的温度,T=1.03–1.94)、jev-lite(在选项字母位置读标签 token)、Open-Jev(由 Yes-minus-No 读出初始化的标量头,T=1.897)、jev-smol(CLEF 式逐选项 logit 打分)
成对打分(交叉编码器)Pair scoring
怎么做:把 (state, question, option) 拼成一个序列整体过一遍编码器,输出一个分数;本质是把 reranker 的结构搬到类型化决策上。
优点:候选与上下文可以互相注意,交互建模最充分;常从检索 reranker 初始化,冷启动快。
代价:每个选项都要过一次前向,选项多时算力线性增长;比前几种慢。
代表实现:open-jev-deberta(三层打分头,池化后的问句与选项表示)、System One scorer(序列分类头逐三元组打分)、GLiNER2.5-Decide(编码器同时做分类、抽 span 与关系)
受限解码Constrained decoding
怎么做:仍然逐 token 生成,但用语法/状态机把输出约束在合法答案集合内,本质是「生成 + 解析」的自动化版本。
优点:改动最小,可以套用在任何现成 LLM 上;能顺带产出简短理由。
代价:没有摆脱自回归,延迟与成本与普通 LLM 同级;概率不是模型分布的直接读出,校准通常最差。
代表实现:djev / razorback16 diffgemma(vLLM 上的 DiffusionGemma 实验路径)、各类把 LLM 变成决策模型的社区适配层

三条训练路径

冻结主干 + 只训 head 最低(笔记本级)
把大模型主干冻住,只在最后一层隐状态上加一个小 head,甚至把候选特征缓存下来只训 head。minojev 用这种方式在 Apple Silicon 笔记本上 37 分钟、峰值约 4 GB 完成一轮;Cloudflare Clef 也是冻结 Qwen 主干、只训 joint schema head 与 rank-256 适配器。
取舍:最省算力,但性能天花板由冻结主干决定;主干没内化的知识,head 学不出来。
实例:minojev、Cloudflare Clef / Clef-flash
LoRA / 全参微调主干 + 读出头 中(单卡可做)
在主干上加 LoRA 或直接全参微调,同时训一个读出结构(标签 logit、pointer head 或标量头)。这是目前开权重模型的主流做法:kev 用 LoRA + pointer head,Open-Jev 用 LoRA rank 8 + 标量头,decider 用 merged LoRA rank 64,Tev1 用 rank 8 / 1 epoch / lr 5e-5 并在 4B 上公开了完整配方。
取舍:性价比最好,能改变模型愿意读出什么;但要小心灾难性遗忘与训练/推理的 prompt 对齐。
实例:kev、Open-Jev、decider、Tev1、JevK5、jev-smol
编码器 + head(从零或从 reranker 初始化) 低(可端侧)
不走解码器路线,直接用 ModernBERT / DeBERTa 这类编码器,从检索 reranker 初始化再训一个打分头。bekko 用 Ettin-reranker 初始化,17M 版本导出 ONNX+INT8 只有 29 MB,能在浏览器 CPU 上跑;laya 用 ModernBERT-large,在 M3 Max 上短决策只要 5–14 毫秒。
取舍:极快极省,端侧友好;但泛化能力明显弱于大主干(S1MB 泛化子集上 bekko-400m 只有 54 分,Jev 是 96 分)。
实例:bekko-system-one-v0、laya、von、open-jev-deberta

校准与目标函数

决策模型的核心卖点是概率能拿去比阈值,所以校准不是加分项而是必选项。一个不校准的 0.9 会让你的阈值策略整体失效。

温度拟合(temperature scaling)
最简单有效:在留出集上拟合一个标量温度,除以 logits 后再 softmax。kev 逐 checkpoint 拟合(T=2.35 / 2.30),decider 按答案类型分别拟合,Bespoke Nimble 用 T=2.179,Open-Jev 用 T=1.897。
label-smoothed CE + Brier loss
Cloudflare Clef 的做法:交叉熵负责把答案选对,Brier loss 负责把概率压准。两者相加是准确率与校准的常见组合。
RLCD(Reinforcement Learning for Calibrated Decisions)
TypeSafe 用于 Jev 的方法:给相邻的有序选项部分分、奖励完全正确的整条记录输出,并加一个 reference penalty 防分布漂移。目前只有博客级别的描述,没有公开实现。
评测指标
至少同时看三个:accuracy(选对没有)、Brier(概率准不准)、ECE 与 over95(高置信度时的错误率)。Decision Index 对每个模型都报了这三类。

各家族的读出方式与校准(第三方普查)

家族读出方式校准
Laya两层 option-marker 打分器 + act/escalate 头 按问题类型与选项数各拟合一个温度;原始 ECE 0.466 → 0.081
KevLoRA 适配器 + 读选项 logit 的 pointer head 逐 checkpoint 拟合 T=2.35(0.8B)/ 2.30(9B)
Bespoke NimbleLoRA 适配器,直接给允许答案的 token 打分 T=2.179;后续版本改用默认温度
decider从 LM head 读字母 logit,再除以存好的温度 按答案类型 1.03–1.94
OpenThai-SystemOne256 路槽位头替换 LM head 按类型学习温度(choice 1.055 / noul 1.047 / score 1.008)
vonoption-marker 头 + 顺序无关注意力掩码 输入条件化校准映射;ECE 0.045–0.109
Decision-1.0共享候选头,读候选端点与 query 向量 未披露
Open-Jev由 Yes-minus-No 读出初始化的标量头 512 条校准集拟合 T=1.897;测试集 ECE 0.0077
open-jev-deberta池化问句与选项表示上的三层打分头 验证集上做事后温度缩放;域内 ECE 0.022
System One scorer对每个 (state, question, option) 三元组的序列分类头 T=1.75;ECE 0.135 → 0.044
jev-lite在选项字母位置读标签 token 见模型卡
SemIf从冻结主干读出选项概率 按工作负载逐次温度校准
Intern-Decision语言主干微调 + 每字段一个 decision token 随模型发布
Cloudflare Clef冻结主干 + joint schema head,并行给所有选项打分 label-smoothed CE + Brier loss,再加 RLCD;ECE 0.040
jev-smolCLEF 式逐选项 logit 打分,语言塔 LoRA 2 epoch 后 Brier 0.610 → 0.457

自己后训练:六步

1. 先确认值不值得训
如果通用模型在你的数据上已经达标,就别训。后训练的价值来自「通用模型做不好、而你会反复做同一个判断」的场景。
2. 造数据
两条路:把自己历史上的判断记录整理成 (state, question, 正确答案);或者用最强的模型在你的数据上跑一遍,人工抽检后作为标签。数据量上,Tev1 用了 37,840 条训练样本,minojev 用了 8,000 条,jev-smol 靠 5 个确定性生成器做到无限扩充。数据量比 epoch 数重要。
3. 选基座与路径
要便宜快:编码器 + head(bekko、laya 路线)。要质量:LoRA 微调 4B–27B 主干(kev、Open-Jev、decider 路线)。要多模态:SmolVLM2-500M 或 Intern-Decision 路线,冻结视觉塔、只对语言塔做 LoRA。
4. 决定读出方式
推荐 per-option logit 打分(标签 logit 读出),比「生成 JSON 再解析」更稳、更快、也更好校准,因为概率直接来自模型分布。
5. 校准单独做一步
在留出集上拟合温度,用 accuracy + Brier + ECE 三个指标验收。别忘了检查置信度超过 95% 时的错误率(over95)——那是阈值策略最容易翻车的地方。
6. 接评测对打
用 S1MB 的 adapter 协议(predict / metadata / close,概率必须对声明选项归一)接入,就能和公开榜单直接比较。注意 S1MB 只覆盖英文文本,你的业务若含图像或中文,需要自建评测。

它和 LLM 的区别

维度决策模型(System One)大语言模型(LLM)
输出类型化答案 + 概率分布,无自由文本自由文本 / 工具调用
确定性高(同样的输入与选项顺序给同样的分布)低,同问可能不同答
延迟一次前向,几十到几百毫秒自回归逐 token 生成,通常更慢
成本按输入 token 计费,输出免费($0.02–$0.24/百万)输入 + 输出都计费
解析不需要解析,直接分支需要解析或依赖 tool call
可校准核心卖点:概率可拿去比阈值通常不校准,置信度不可靠
适合分类、路由、审核、打分、门控生成、推理、写代码、多轮对话

统一 API 现状

词汇已经统一(noul / choice / score,OpenAI 叫 predicate / choice / score), 但端点仍然分裂:目前没有任何一个 API 能同时调用多家厂商的多模态决策模型。 OpenRouter 是唯一结构上跨厂商的决策端点,但它当前文档只覆盖了 TypeSafe 的 Jev,且 Jev 只支持文本。

平台端点多厂商图像输入原语价格微调链接
OpenAI Decisions API
OpenAI
POST https://api.openai.com/v1/decisions 单一厂商 支持 predicate / choice / score $0.10 / 百万输入 token(输出免费) 不支持 docs ↗
playground ↗
Perplexity Decisions API
Perplexity
POST https://api.perplexity.ai/v1/decisions 单一厂商 支持 noul / choice / score $0.02 / 百万输入 token(输出免费) 不支持 docs ↗
api_ref ↗
weights ↗
OpenRouter Decisions 路由
OpenRouter
POST https://openrouter.ai/api/alpha/decisions · POST https://openrouter.ai/api/v1/systemone 跨厂商 未定义 noul / choice / score 按输入 token 计费(各模型不同),输出 token 一律免费 不支持 rankings ↗
models_api ↗
jev_docs ↗
api_ref ↗
TypeSafe System One API (Jev)
TypeSafe AI
POST https://api.typesafe.ai/v1/systemone 单一厂商 不支持 choice / noul / score $0.042 / 百万输入 token(输出免费) 不支持 docs ↗
models ↗
api ↗
launch ↗
sdk_py ↗
sdk_js ↗
Cloudflare Workers AI (Clef)
Cloudflare
POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef 单一厂商 支持 noul / choice / score $0.24/百万(Clef)· $0.09/百万(Clef-flash) 支持 docs ↗
blog ↗
weights ↗
rl_interest ↗
AWS Bedrock Converse
AWS
bedrock-runtime Converse / ConverseStream 跨厂商 支持 — 按模型计费 支持 docs ↗
custom_models ↗
Azure AI Foundry
Microsoft
托管计算 / serverless REST 跨厂商 支持 — 按模型计费 支持 docs ↗
vision_ft ↗
Hugging Face Inference Providers
Hugging Face
HF Router 跨厂商 未定义 — 按 provider 计费 不支持 docs ↗
Together AI 视觉微调
Together AI
托管微调 + OpenAI 兼容推理 单一厂商 支持 — 按处理 token 计费 + 按分钟托管费 支持 vision_ft ↗
models ↗
Fireworks AI 视觉微调
Fireworks AI
托管训练 + 专用部署 单一厂商 支持 — 按训练与部署计费 支持 finetune ↗
vlm ↗
自托管开源决策模型
社区 / 多家
本地推理(多数兼容 Jev 的 /v1/systemone 线格式) 跨厂商 支持 noul / choice / score 无 API 费用,只有自己的 GPU 成本 支持 s1mb ↗
awesome ↗
jew_smol ↗

术语表(36 条)

Calibrated Probability
A calibrated probability is a number from a model that means what it says: across all the answers given at 80%, about 80% are right. Jev returns one for every question, a distribution over the options for Choice and Score and a single yes probability for Noul.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Calibrated Probability Calibrated Probability A calibrated probability is a number from a model that means what it says: across all the answers given at 80%, about 80% are right. Jev returns one for every question, a distribution over the options for Choice and Score and a single yes probability for Noul. Full guide: Choice, Score and Noul: the three primitives The distinction worth holding onto is between a number that ranks and a number that predicts. A softmax output from an ordinary classifier ranks options fine but often runs hot, clustering near 1.0 whether or not the model deserves it. A calibrated probability is meant to be read as a frequency, so 0.7 means seven times in ten. For a Choice question the probabilities cover the listed options and sum to 1. The docs’ ticket routing example returns {"choice":"returns","confidence":1.0,"probabilities":{"shipping":0.0,"returns":1.0,&qu
Calibration
Calibration is the match between a model's stated probabilities and how often it turns out to be right. A calibrated model that answers 80% across a large batch of questions is correct on about 80% of them. Calibration is about the honesty of the numbers, separate from raw accuracy.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Calibration Calibration Calibration is the match between a model's stated probabilities and how often it turns out to be right. A calibrated model that answers 80% across a large batch of questions is correct on about 80% of them. Calibration is about the honesty of the numbers, separate from raw accuracy. Full guide: RLCD explained: Reinforcement Learning for Calibrated Decisions Accuracy and calibration come apart. A model can be right 95% of the time and still be badly calibrated if it says 99% every time, and a model that is right only 60% of the time can be perfectly calibrated if it says 60%. The second kind is more useful to software, because a threshold in your code does what you expect. The docs’ bug severity example shows the shape. A ticket about the export button crashing in Safari gets levels 0 to 2, and the answer comes back with probabilities of 0.0, 0.7 and 0.3, a score of 1.3, and a confidence of 0.54. If the model is calibra
Choice
A Choice is a System One question type for selecting one option from a defined set. The answer names the highest-probability option, includes the full distribution over the options, which sums to 1, and a confidence value from 0 to 1. A Choice question accepts up to 255 options.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Choice Choice A Choice is a System One question type for selecting one option from a defined set. The answer names the highest-probability option, includes the full distribution over the options, which sums to 1, and a confidence value from 0 to 1. A Choice question accepts up to 255 options. Full guide: Choice, Score and Noul: the three primitives The docs’ example routes a support ticket to a department. With the options shipping, returns and billing, the answer comes back as {"type":"choice","choice":"returns","confidence":1.0,"probabilities":{"shipping":0.0,"returns":1.0,"billing":0.0}}. Your code reads choice to route and reads confidence to decide whether to route at all. One line of advice in the docs is easy to skip and costly when you do. When your list of options might not cover every real case, add an “other” or “none of the above” option. Without
Composite scoring
Composite scoring is one of TypeSafe's documented patterns, described as "Combine several dimensions of analysis into a single score." You ask a separate question about each dimension, then your own code weights and combines the answers. The model never picks the weights, and the arithmetic stays deterministic.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Composite scoring Composite scoring Composite scoring is one of TypeSafe's documented patterns, described as "Combine several dimensions of analysis into a single score." You ask a separate question about each dimension, then your own code weights and combines the answers. The model never picks the weights, and the arithmetic stays deterministic. Full guide: How to build with System One models TypeSafe’s build guidance separates the two halves. Ask many independent parallel questions about the same state, then “compose answers with deterministic rules or code-controlled weights”, keeping deterministic work in code. That split matters here, because arithmetic and counting are on jev-1.13’s published list of weak spots. You do not want the model doing the summing. A lead-scoring example. Ask four Score questions against the same state, each with its own ordered levels: budget fit, urgency, decision-making authority, technical fit. Eac
Confidence Gating
Confidence gating is the practice of branching your code on a model's confidence value as well as on its answer. High confidence runs the action automatically. Low confidence sends the case to a person or a fallback. TypeSafe calls the pattern confidence-gated routing and leaves the thresholds to you.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Confidence Gating Confidence Gating Confidence gating is the practice of branching your code on a model's confidence value as well as on its answer. High confidence runs the action automatically. Low confidence sends the case to a person or a fallback. TypeSafe calls the pattern confidence-gated routing and leaves the thresholds to you. Full guide: How to build with System One models A classifier that only returns a label forces one behaviour for every prediction. Adding confidence gives a second axis, so the same answer can trigger an automatic action in one case and a review queue in another. The docs put numbers on it. Anything under 0.5 counts as uncertain and goes to a human. Destructive operations, their example is a funds transfer, need confidence above 0.9 before they run without a confirmation step, while a read-only balance check can proceed at a lower bar. The thresholds scale with what happens if the answer is wrong, and the docs
Confidence
Confidence is a number from 0 to 1 that TypeSafe computes from the probability distribution an answer already carries. A concentrated distribution gives high confidence, a spread one gives low. Choice and Score answers include it. Noul does not, because its single probability already carries the certainty.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Confidence Confidence Confidence is a number from 0 to 1 that TypeSafe computes from the probability distribution an answer already carries. A concentrated distribution gives high confidence, a spread one gives low. Choice and Score answers include it. Noul does not, because its single probability already carries the certainty. Full guide: Choice, Score and Noul: the three primitives Nothing extra is being measured here. The docs describe confidence as “a statistic computed from the probability distribution the answer already gives you”, collapsing the shape of that distribution into one number. A Choice that puts 1.0 on returns and 0.0 everywhere else comes back with confidence 1.0. A Score split 0.7 and 0.3 across two adjacent levels comes back with confidence 0.54. TypeSafe’s guidance runs in three bands. High confidence means act automatically. Medium means proceed with caution, asking the user to confirm or flagging the case for review. Low
Decision Model
Decision model is the generic label people reach for when describing an AI model whose output is a typed answer with a probability attached rather than prose. It is not an official term from any vendor. System One model is TypeSafe's name for the same idea, and Jev is the first shipping example.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Decision Model Decision Model Decision model is the generic label people reach for when describing an AI model whose output is a typed answer with a probability attached rather than prose. It is not an official term from any vendor. System One model is TypeSafe's name for the same idea, and Jev is the first shipping example. Full guide: What is a System One (System 1) model? The phrase gets used two ways. Loosely, it covers any model that hands software a decision rather than a paragraph, which includes encoder classifiers like BERT and DeBERTa, span taggers like GLiNER, and anything wrapped in constrained decoding to force a fixed output shape. More narrowly, people use it for the new category TypeSafe is trying to define, where the typed answer and its probability are what the model was trained to produce in the first place. There is no standards body here. Treating “decision model” as a description rather than a spec avoids arguing about w
Expected calibration error
Expected calibration error (ECE) is the average gap between a model's stated confidence and how often it is actually right. You bucket predictions into bins by confidence, compare each bin's mean confidence to its accuracy, and average the differences weighted by how many predictions fall in each bin.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Expected calibration error Expected calibration error Expected calibration error (ECE) is the average gap between a model's stated confidence and how often it is actually right. You bucket predictions into bins by confidence, compare each bin's mean confidence to its accuracy, and average the differences weighted by how many predictions fall in each bin. Full guide: RLCD explained: Reinforcement Learning for Calibrated Decisions A worked version. Take every prediction a model made with confidence between 0.80 and 0.90, where the mean confidence in that bin works out to roughly 0.85. If the model was right on 62% of them, the bin’s gap is about 0.23. Repeat for each bin, weight by bin size, and the total is the ECE. A perfectly calibrated model scores 0, lower is better, and the figure only means something alongside the bin count and the dataset it was measured on. The measure is standard in the calibration literature. The widely cited ref
Hallucination
Hallucination is when a language model states something false in generated text. A System One model never writes free text, so it cannot invent a fact in prose. It can still be wrong: it can pick the wrong option or return a probability that does not match how often it is right.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Hallucination Hallucination Hallucination is when a language model states something false in generated text. A System One model never writes free text, so it cannot invent a fact in prose. It can still be wrong: it can pick the wrong option or return a probability that does not match how often it is right. Full guide: System One models vs LLMs The usual failure of a chat model is fluent invention: a citation that does not exist, an API method that was never shipped. Jev returns a typed answer drawn from options you defined, plus a probability distribution over them. There is no prose for a fabricated fact to hide in. That removes one failure mode. It does not remove error. TypeSafe’s own jaggedness page for jev-1.13 lists nine things the model does badly, including reading dates as text rather than as ordered quantities, and no guarantee that P(A) and P(not A) sum to 1 across related questions. A Choice that comes back wrong at high confidence is
Intent routing
Intent routing is one of TypeSafe's documented patterns: "Classify a user's intent and route to the appropriate handler." A Choice question returns the intent plus a probability for every option, and your code maps the winning option to a handler. The model does not call the handler.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Intent routing Intent routing Intent routing is one of TypeSafe's documented patterns: "Classify a user's intent and route to the appropriate handler." A Choice question returns the intent plus a probability for every option, and your code maps the winning option to a handler. The model does not call the handler. Full guide: How to build with System One models This is the pattern most people reach for first, and it shows the split between model and code clearly. TypeSafe’s framing: “System One is TypeSafe’s model for building AI-powered software, not agents. It does not generate code or choose its own next action.” The Choice answer is data. Your router is a switch statement. From the docs, a Choice answer for a support ticket looks like {"type":"choice","choice":"returns","confidence":1.0,"probabilities":{"returns":1.0,"shipping":0.0,"billing&q
Jaggedness
Jaggedness is TypeSafe's word for the published, per-version list of what a model does badly. Each Jev version gets its own page. The one for jev-1.13 documents nine failure modes, so you can design around known weak spots instead of discovering them in production.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Jaggedness Jaggedness Jaggedness is TypeSafe's word for the published, per-version list of what a model does badly. Each Jev version gets its own page. The one for jev-1.13 documents nine failure modes, so you can design around known weak spots instead of discovering them in production. Full guide: What is Jev AI? TypeSafe's decision model explained Publishing a per-version failure list is unusual, and it helps because the entries are specific. Date and time comparison is one of them: “Jev reads dates as text, not as ordered quantities”, so asking whether one timestamp comes before another is unreliable. Math and numbers is another, with arithmetic and counting called out directly. Adversarial content is a third worth planning around, because “State is data, and jev-1.13 does not treat it as hostile by default”. The fourth is structural: there is no guarantee that P(A) and P(not A) sum to 1 across related questions, so two questions that l
Jev
Jev is the first System One model, announced by TypeSafe AI on 2026-09-15. It answers typed questions about a block of text and returns a choice, a score, or the probability that a yes/no statement is true. Access is a closed managed API, open to anyone since TypeSafe removed the waitlist on 2026-09-20. Weights have not been released.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Jev Jev Jev is the first System One model, announced by TypeSafe AI on 2026-09-15. It answers typed questions about a block of text and returns a choice, a score, or the probability that a yes/no statement is true. Access is a closed managed API, open to anyone since TypeSafe removed the waitlist on 2026-09-20. Weights have not been released. Full guide: What is Jev AI? TypeSafe's decision model explained A request sends one state, the text to evaluate, plus a set of named questions, and comes back with a typed answer under each name. All the questions in a call run against the same state, so asking four things about one ticket costs one round trip to POST https://api.typesafe.ai/v1/systemone. The shipping version is jev-1.13.0. The aliases jev-latest and jev-preview both point at it. Input costs $0.042 per million tokens and output is free. A request carries 64k tokens in total, with 32k as the ceiling on the state plus the longest single qu
Non-autoregressive
Non-autoregressive describes a model that does not build its output one token at a time. Autoregressive models predict each token conditioned on the tokens before it. TypeSafe says Jev generates all outputs in a single query instead, though the architecture behind that claim has not been published.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Non-autoregressive Non-autoregressive Non-autoregressive describes a model that does not build its output one token at a time. Autoregressive models predict each token conditioned on the tokens before it. TypeSafe says Jev generates all outputs in a single query instead, though the architecture behind that claim has not been published. Full guide: Jev architecture: parameters, encoder or decoder, and no paper yet Start with the thing it contrasts against. An autoregressive language model produces text by predicting one token, feeding that token back in, then predicting the next. Each step waits for the one before it, which is why a long answer takes longer to return than a short one. TypeSafe’s launch post describes Jev’s sampling as “Parallel”, which “Generates all outputs in a single query” versus sequential token generation, and calls it “Incredibly efficient and hardware-aware”. That is the whole public claim. Parameter count, base architectu
Noul
A Noul is a question type in System One AI models such as Jev: it asks a yes/no question and returns one number, the probability that the answer is yes, from 0 to 1. It carries no separate confidence value, because the probability is itself the certainty measure. The docs do not explain the name.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Noul Noul A Noul is a question type in System One AI models such as Jev: it asks a yes/no question and returns one number, the probability that the answer is yes, from 0 to 1. It carries no separate confidence value, because the probability is itself the certainty measure. The docs do not explain the name. Full guide: Choice, Score and Noul: the three primitives In practice a Noul is a named question in the request with its type set to noul. The docs’ example sends a support message as the state and asks {"type":"noul","instructions":"Is the customer asking for a human agent?"}. The answer comes back as {"type":"noul","noul":0.99}: a 99% probability that the answer is yes. There is no special Noul data type in the response. The value is a plain number, and your code compares it against a threshold like any other float. Reading the number is direct. A value near 1 is a strong ye
Question
A question is the typed object you attach to state in a System One request. Each one names what you want decided and the shape of the answer. Jev supports three question types: Choice for picking one option, Score for rating against ordered levels, and Noul for a yes/no probability.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Question Question A question is the typed object you attach to state in a System One request. Each one names what you want decided and the shape of the answer. Jev supports three question types: Choice for picking one option, Score for rating against ordered levels, and Noul for a yes/no probability. Full guide: Choice, Score and Noul: the three primitives Questions are keyed in a map, so each answer comes back under the name you gave it. Choice selects one option from a defined set and returns the winning option, the full probability distribution and a confidence value, with up to 255 options allowed. Score evaluates content against ordered, descriptive levels, minimum 2 and maximum 10, and returns a probability-weighted mean of the level numbers. Noul returns a single number between 0 and 1, the probability that a yes/no statement is true, and no separate confidence value. The docs advise asking “the most explicit, narrow, specific, atomic ques
RLCD
RLCD stands for Reinforcement Learning for Calibrated Decisions, the training method TypeSafe AI says it used for Jev. Per the launch post it optimizes for answers with epistemically honest probabilities on System One tasks, rather than for text a human rater prefers or output a program can check.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary RLCD RLCD RLCD stands for Reinforcement Learning for Calibrated Decisions, the training method TypeSafe AI says it used for Jev. Per the launch post it optimizes for answers with epistemically honest probabilities on System One tasks, rather than for text a human rater prefers or output a program can check. Full guide: RLCD explained: Reinforcement Learning for Calibrated Decisions The name sits alongside two older acronyms. RLHF, reinforcement learning from human feedback, trains a model toward responses human raters prefer. RLVR, reinforcement learning from verifiable rewards, trains it toward outputs a program can check as correct. RLCD swaps the target again: the reward is tied to whether the stated probability matches how often the answer turns out to be right. That target is what calibration means. A model trained this way should be right on roughly 80% of the answers it labels 80%, and wrong on roughly 20% of them, which is what makes a th
Score
A Score is a System One question type that rates content against ordered, descriptive levels, minimum 2 and maximum 10. The returned score is the probability-weighted mean of the level numbers, so it can land between levels. The answer also carries a legend, the level probabilities and a confidence value.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Score Score A Score is a System One question type that rates content against ordered, descriptive levels, minimum 2 and maximum 10. The returned score is the probability-weighted mean of the level numbers, so it can land between levels. The answer also carries a legend, the level probabilities and a confidence value. Full guide: Choice, Score and Noul: the three primitives Levels are written as short descriptions, which is what lets the model place content without a separate rubric hidden in your prompt. The docs’ bug severity example uses three: “Cosmetic; no impact to functionality”, “Broken or degraded feature, but workaround exists”, and “Blocking issue; no workaround exists”. Given the state “The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari”, the answer is a score of 1.3 with probabilities of 0.0, 0.7 and 0.3 on levels 0, 1 and 2, at confidence 0.54. The arithmetic is plain
Speculative fan-out
Speculative fan-out is one of TypeSafe's documented patterns: "Send many questions in a single call, including speculative ones, and let your code decide what's relevant." Because questions are evaluated in parallel, extra questions typically cost no extra latency, so you ask ahead instead of making a second round trip.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Speculative fan-out Speculative fan-out Speculative fan-out is one of TypeSafe's documented patterns: "Send many questions in a single call, including speculative ones, and let your code decide what's relevant." Because questions are evaluated in parallel, extra questions typically cost no extra latency, so you ask ahead instead of making a second round trip. Full guide: How to build with System One models The docs state the mechanism directly: “All questions are evaluated in parallel, so adding more questions to a call typically doesn’t add any latency to the response.” That flips the usual cost model. With a chat model you ask for the minimum, because every extra field is more tokens to generate and more time to wait. Here the marginal question is close to free in latency terms, so you ask everything any downstream branch might plausibly need, then throw away what the branch you took did not use. Concrete shape. One call on a
State
State is TypeSafe's word for the content you hand a System One model: the ticket, document, transcript or JSON object the questions are asked about. It is passed as text, and the state plus the longest question must fit within 32k tokens on Jev.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary State State State is TypeSafe's word for the content you hand a System One model: the ticket, document, transcript or JSON object the questions are asked about. It is passed as text, and the state plus the longest question must fit within 32k tokens on Jev. Full guide: How to build with System One models State is the data half of a System One request. The other half is the set of typed questions you attach to it. Jev accepts text only, passed as a string or as structured text (a JSON object or an array of text values). No image, audio or video input. Total context is 64k tokens per request, and state plus the longest single question is capped at 32k tokens. TypeSafe’s build guidance is blunt about what belongs in there: “Include only the context relevant to the current questions. This helps the model avoid distractions and context rot.” That advice has teeth. The published jaggedness page for jev-1.13 lists large state full of irrelevant deta
Structured outputs
Structured outputs is the existing technique for making a language model emit valid JSON: constrain decoding to a schema so every generated token keeps the output parseable. OpenAI and others ship it in their APIs, and libraries such as Outlines and Instructor do the same over open models.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Structured outputs Structured outputs Structured outputs is the existing technique for making a language model emit valid JSON: constrain decoding to a schema so every generated token keeps the output parseable. OpenAI and others ship it in their APIs, and libraries such as Outlines and Instructor do the same over open models. Full guide: System One models vs LLMs Mechanically, schema-constrained decoding masks the logits at each step so only tokens the grammar allows can be sampled. The model still generates a token at a time. A JSON object with six fields costs six fields’ worth of sequential decoding, and any confidence figure has to be reconstructed from logprobs afterwards. A System One model differs at exactly that step. Nothing is generated and then constrained, because the option set you passed in is the answer space, and a calibrated probability over those options comes back as part of the answer. TypeSafe’s public claim is that Jev “Gen
System One Model
A System One model is a class of AI models built to make fast, structured decisions that software can use directly. It returns typed decisions and probabilities instead of generated text, so it does not write replies, produce code, or explain its reasoning. Jev by TypeSafe AI is the first one.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary System One Model System One Model A System One model is a class of AI models built to make fast, structured decisions that software can use directly. It returns typed decisions and probabilities instead of generated text, so it does not write replies, produce code, or explain its reasoning. Jev by TypeSafe AI is the first one. Full guide: What is a System One (System 1) model? The category name comes from Daniel Kahneman’s Thinking, Fast and Slow, which splits thinking into fast, intuitive System 1 and slow, deliberate System 2 reasoning. A System One model takes the fast half: a judgment your code needs in milliseconds, returned in a shape your code can branch on without parsing prose. TypeSafe’s docs define three question types for it, Choice, Score and Noul, and the model never picks the next action, so your code keeps the control flow and every side effect. This entry is the short definition. The full explanation, with the three question type
System Two Model
System Two model is a contrast term, not a product. No vendor ships a model branded this way. The phrase points to the slow, deliberate half of Daniel Kahneman's split in Thinking, Fast and Slow, which in practice means a reasoning LLM that works through text before it answers.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary System Two Model System Two Model System Two model is a contrast term, not a product. No vendor ships a model branded this way. The phrase points to the slow, deliberate half of Daniel Kahneman's split in Thinking, Fast and Slow, which in practice means a reasoning LLM that works through text before it answers. Full guide: System One models vs LLMs The term only exists because TypeSafe named its own category after Kahneman’s System 1. Its launch post cites “fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning” as the source of the name. Once one half has a product label, people reach for the other half to describe everything that is not a System One model: chat models and reasoning models, anything that works in tokens and hands you prose. Nobody sells a “System Two model”. If you see the phrase in a comparison table, read it as shorthand for a reasoning LLM. The practical difference is time budget. TypeSafe claims Jev a
Thinking, Fast and Slow
Thinking, Fast and Slow is Daniel Kahneman's 2011 book (Farrar, Straus and Giroux) describing two modes of human thought: a fast, automatic System 1 and a slow, deliberate System 2. TypeSafe borrowed the framing for the name "System One model", the category Jev launched under in September 2026.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Thinking, Fast and Slow Thinking, Fast and Slow Thinking, Fast and Slow is Daniel Kahneman's 2011 book (Farrar, Straus and Giroux) describing two modes of human thought: a fast, automatic System 1 and a slow, deliberate System 2. TypeSafe borrowed the framing for the name "System One model", the category Jev launched under in September 2026. Full guide: What is a System One (System 1) model? The launch post makes the borrowing explicit, contrasting “fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning” and placing the new model class on the fast side: quick typed decisions that software consumes, with no written reasoning attached. Kahneman’s subject was people. System 1 in the book describes human cognition, its speed and its characteristic biases, and makes no claim about how any machine is built. Using the name for a class of models is an analogy about the job being done, not evidence about architecture. What
Typed output
Typed output is a model answer that arrives already shaped as a value your code can use, such as an enum member or a probability between 0 and 1. There is no parsing step and no schema validation on a text blob. A System One model returns typed output directly rather than describing an answer in prose.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Typed output Typed output Typed output is a model answer that arrives already shaped as a value your code can use, such as an enum member or a probability between 0 and 1. There is no parsing step and no schema validation on a text blob. A System One model returns typed output directly rather than describing an answer in prose. Full guide: System One models vs LLMs TypeSafe describes System One models as returning “typed answers and probabilities rather than generated text”, and says they “do not write replies, produce code, or generate explanations of their reasoning.” What comes back from a Choice question is concrete. The docs show {"type":"choice","choice":"returns","confidence":1.0,"probabilities":{"returns":1.0,"shipping":0.0,"billing":0.0}}. In the JavaScript SDK you read it as response.answers.category.choice, and the value is one of the keys you sup
TypeSafe AI
TypeSafe AI is the San Francisco company that coined the term System One model and shipped the first one, Jev, on 2026-09-15. It was founded by Diogo Almeida (CEO), Erik Gafni (CTO) and Sasha Sheng (COO), and raised a $40M seed round led by DCVC.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary TypeSafe AI TypeSafe AI TypeSafe AI is the San Francisco company that coined the term System One model and shipped the first one, Jev, on 2026-09-15. It was founded by Diogo Almeida (CEO), Erik Gafni (CTO) and Sasha Sheng (COO), and raised a $40M seed round led by DCVC. Full guide: What is Jev AI? TypeSafe's decision model explained The company’s pitch starts from a question Almeida puts at the top of the launch post: “Models have been superhuman at chat for years, so where is all the automation?” Its answer is a model that returns a typed decision your code can branch on, with a probability attached, instead of a paragraph you have to parse. TypeSafe’s own team page says Almeida co-invented RLHF and InstructGPT and previously worked at Google Brain, Gafni is a repeat founder with production AI experience, and Sheng was a research engineer at Meta’s FAIR. That background matters mainly because the training method behind Jev, RLCD, is describe
Zero-shot classifier
A zero-shot classifier assigns text to labels it was never trained on, using the label descriptions you supply at call time rather than labelled examples. Jev's Choice question works this way: you define the categories in the request, and the model returns a probability distribution over them.
展开
Skip to content System OneSystem One ModelsModelsUse casesExamplesRecipesGuidesIntegrationsGlossary Search⌘K Submit Menu Site navigation ModelsUse casesExamplesRecipesGuidesIntegrationsGlossarySubmit an example Close Home Glossary Zero-shot classifier Zero-shot classifier A zero-shot classifier assigns text to labels it was never trained on, using the label descriptions you supply at call time rather than labelled examples. Jev's Choice question works this way: you define the categories in the request, and the model returns a probability distribution over them. Full guide: Is Jev just a zero-shot classifier? The term predates Jev by years. Encoder classifiers such as BERT and DeBERTa, and span models such as GLiNER, have been doing classification against labels supplied at inference time for a while, which is why the comparison came up straight after the launch on 2026-09-15. agentpedia.codes reports that a Hacker News commenter called Jev “basically a zero-shot classifier” and that Diogo Almeida replied “exactly right!”. That exchange could not be confirmed in the Hacker News thread itself, so treat it as agentpedia.codes’ account rather than a direct citation from HN
System One 模型(System 1)
返回类型化决策与校准概率、而不是生成文本的一类 AI 模型。
展开
名字借自双过程理论:System 1 是快速直觉,System 2 是慢速推理。System One 模型承担前者——在流程里做快速、结构化、可直接被代码消费的判定。第一个是 TypeSafe AI 的 Jev(2026-09-15)。
RLCD
Reinforcement Learning for Calibrated Decisions,TypeSafe 用于 Jev 的训练方法。
展开
公开信息只到博客级别:给相邻的有序选项部分分、奖励完全正确的整条记录输出、并加一个 reference penalty 防止分布漂移。具体损失与数据未公开,也没有公开实现。
Brier 分数
衡量概率预测准确度的严格评分规则,越低越好。
展开
对二分类是 (p - y)² 的均值。S1MB 的 Noul 任务用它作为主指标,并用 balanced accuracy 做基线校正后折算成 0–100 的任务分。
ECE(期望校准误差)
把预测按置信度分桶,比较每桶的平均置信度与实际准确率。
展开
ECE 越小说明「说 90% 就是 90%」。决策模型的价值很大程度依赖这个性质,因为代码要拿概率去比阈值。Decision Index 对每个模型报告 ECE、Brier 与 over95(置信度 >95% 的错误率)。
per-option logit 打分
不生成文本,而是把每个选项的标签放在答案位置做 teacher forcing,取平均 token log-prob 再 softmax。
展开
这是当前开源决策模型最主流的实现方式(jev-smol、decider、JevK5/SemIf、Intern-Decision 等)。相比「生成 JSON 再解析」,它更快、更稳、也更好校准,因为概率直接来自模型分布。
joint schema head
在冻结 backbone 的隐状态上加一个小 transformer head,联合给所有问题的所有选项打分。
展开
Cloudflare Clef 的架构:backbone 只做一次 prefill,head 负责把证据路由到每个问题、并让字段之间互相关注,最后一次性输出所有选项的 logit。
confidence 与概率的区别
probability 是每个选项的分布,confidence 是模型对「这次选择」的把握。
展开
Choice 与 Score 会同时返回两者;Noul 在 Jev 上只返回概率、不返回独立 confidence。生产系统通常对 confidence 设阈值,低置信度转人工或升级到更强的模型。
基线校正分(S1MB Task Avg)
不是准确率,而是「相对忽略输入的基线补上了多少差距」。
展开
50 分意味着补上了一半差距。Noul 用 balanced accuracy 把 50% 映射到 0、100% 映射到 100;Choice 取「均匀随机」与「固定作答」中较优者为基线;Score 用 MAE 对常数预测做校正。
Borda Score
把每个基准上的相对名次折算成分数的排序方法。
展开
每个基准第一名 100 分、最后一名 0 分,中间等距;再对所有基准等权平均。它衡量「稳定地名列前茅」,会随参赛模型名单变化而变化,不能与 Task Avg 混用。
决策原语(primitive)
Choice / Noul / Score 三种问法,对应分类、判定与打分。
展开
同一份 state 可以挂多个问题、不同类型混用,一次前向返回全部答案。OpenAI 把它叫 predicate / choice / score,Perplexity 与 S1MB 直接叫 noul / choice / score。

指南

Choice, Score and Noul: the three primitives
Choice, Score and Noul are the three question types a System One model answers. Concrete examples, their limits, and how confidence is derived.
How to build with System One models
State in, typed questions out. Code owns the thresholds and side effects. The four official patterns, and the jaggedness list as design constraints.
How to get Jev API access
How to get Jev API access: open sign-up at the TypeSafe console, the key page, the TYPESAFE_API_KEY env var, SDK installs, and a first working call.
Is Jev just a zero-shot classifier?
Engineers called Jev a zero-shot classifier within hours of launch. What the comparison gets right, what it misses, and what is actually new.
Jev alternatives: hosted, open-source and local options compared
The realistic Jev alternatives in September 2026: other hosted APIs, open models you host, community clones and plain classifiers, with the benchmarks behind each.
Jev architecture: parameters, encoder or decoder, and no paper yet
TypeSafe has not published Jev's parameter count, base architecture or weights. Here is the public record and what the community has guessed from it.
Jev benchmarks: every public evaluation in one place
Every public Jev benchmark we could verify: TypeSafe's own evals, independent tests against Opus, Haiku, GPT and Laya, and what they do and don't show.
Jev game demos: Doom, Pokemon, Mario, drones and more
Every Jev game and real-time control demo in the directory: Doom, Pokemon, StarCraft, Mario, Tetris, Snake, chess, poker, drones, robot arms and driving.
Jev speed and pricing
Jev costs $0.042 per million input tokens with free output. TypeSafe claims 70ms to 500ms and 40x to 200x. What those numbers cover.
RLCD explained: Reinforcement Learning for Calibrated Decisions
RLCD is TypeSafe's training method for Jev. What Reinforcement Learning for Calibrated Decisions means, how it differs from RLHF and RLVR, and what is not public.
System One models vs LLMs
System One models return typed decisions; LLMs return text. A side-by-side on output, latency, cost, hallucination and when each one wins.
What is a System One (System 1) model?
A System One (System 1) model is a decision model that returns typed, calibrated answers instead of text. How it works, the name, and which models exist.
What is Jev AI? TypeSafe's decision model explained
Jev is TypeSafe AI's decision model. It answers typed questions about text with calibrated probabilities, and it is not an LLM, because it writes nothing.

横向对比

集成

用例

Agent routing and skill selection with System One models
An agent with a large skill roster picks badly from truncated descriptions. Jev ranks the whole roster in one call and can also answer that nothing fits.
Citation verification with System One models
Check whether a quoted source actually backs the claim built on it. A string match catches fabrications, then one Jev Choice reads the surrounding context.
Compliance verification with System One models
Check a policy or contract against a written rulebook. Jev answers a whole compliance checklist in one typed request and flags the uncertain findings.
Composite scoring with System One models
Break one fuzzy judgment into separate rated dimensions. Jev scores each one, and your code holds the weights, so you can see how a ranking was built.
Confidence-gated actions with System One models
Use Jev's confidence value as a second axis: the answer says what the user wants, confidence says whether your code should act on it without asking.
Entity alignment with System One models
Decide whether two records describe the same thing using one Jev Score whose levels are the outcomes: merge, leave unlinked, or send to a curator.
Feature extraction for machine learning with System One models
Turn free text into numeric columns a gradient-boosting model can train on. Each Jev question becomes a feature, and the probabilities carry the uncertainty.
Hierarchical classification with System One models
Walk a deep taxonomy one level at a time. Each node is a Jev choice question, and the probabilities let you keep several paths open instead of one.
Intent and model routing with System One models
Classify what a request is and how hard it looks with one Jev call, then send it to plain code, a small model, a frontier model or a person.
LLM guardrails with System One models
Screen every message into and out of an LLM app with one Jev call: a probability per hazard, a severity rating, and routing thresholds you own in code.
RAG passage filtering with System One models
Screen every retrieved passage with Jev before it reaches the answering model: keep the usable ones, flag contradictions, drop injected instructions.
Real-time control with System One models
Jev answers fast enough to sit inside a game loop or a UI frame. Your code serialises the current state, asks a typed question, and acts on the answer.
Self-consistency checks with System One models
Ask Jev the same question repeatedly and see how far the answer moves. Measuring that spread tells you where a threshold is safe and where it is not.
Semantic code linting with System One models
Run your team's written conventions as Jev questions in CI. A System One model reads each diff hunk and returns typed judgments your pipeline can act on.
Semantic reranking with System One models
Score each query and candidate pair with one Jev question, then sort the shortlist by that number. TypeSafe's legal retrieval test moved top-1 from 5% to 18%.
Structured extraction with System One models
Jev does not write text, so extraction works in reverse: code finds candidate values, Jev picks the right one, and the value you get back is a verbatim copy.
Support inbox triage with System One models
Triage support tickets in one Jev call: category, severity, refund intent and frustration come back as typed values your routing code can branch on.
Typed tool dispatch with System One models
Pick the tool and fill its arguments with Jev Choice questions over the values each function already accepts, then let your own code make the call.